首页|Can we reconstruct a dysarthric voice with the large speech model Parler TTS?

Can we reconstruct a dysarthric voice with the large speech model Parler TTS?

来源：

英文摘要

Speech disorders can make communication hard or even impossible for those who develop them. Personalised Text-to-Speech is an attractive option as a communication aid. We attempt voice reconstruction using a large speech model, with which we generate an approximation of a dysarthric speaker's voice prior to the onset of their condition. In particular, we investigate whether a state-of-the-art large speech model, Parler TTS, can generate intelligible speech while maintaining speaker identity. We curate a dataset and annotate it with relevant speaker and intelligibility information, and use this to fine-tune the model. Our results show that the model can indeed learn to generate from the distribution of this challenging data, but struggles to control intelligibility and to maintain consistent speaker identity. We propose future directions to improve controllability of this class of model, for the voice reconstruction task.

作者：Ariadna Sanchez、Simon King

作者单位：

学科分类：语言学

推荐引用：Ariadna Sanchez,Simon King.Can we reconstruct a dysarthric voice with the large speech model Parler TTS?[EB/OL].(2025-06-04)[2025-06-28].https://arxiv.org/abs/2506.04397.点此复制

Can we reconstruct a dysarthric voice with the large speech model Parler TTS?

Can we reconstruct a dysarthric voice with the large speech model Parler TTS?

评论