首页|Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding

Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding

来源：

英文摘要

Modern autoregressive speech synthesis models leveraging language models have demonstrated remarkable performance. However, the sequential nature of next token prediction in these models leads to significant latency, hindering their deployment in scenarios where inference speed is critical. In this work, we propose Speech Speculative Decoding (SSD), a novel framework for autoregressive speech synthesis acceleration. Specifically, our method employs a lightweight draft model to generate candidate token sequences, which are subsequently verified in parallel by the target model using the proposed SSD framework. Experimental results demonstrate that SSD achieves a significant speedup of 1.4x compared with conventional autoregressive decoding, while maintaining high fidelity and naturalness. Subjective evaluations further validate the effectiveness of SSD in preserving the perceptual quality of the target model while accelerating inference.

作者：Zijian Lin、Yang Zhang、Yougen Yuan、Yuming Yan、Jinjiang Liu、Zhiyong Wu、Pengfei Hu、Qun Yu

作者单位：

学科分类：声学工程计算技术、计算机技术

推荐引用：Zijian Lin,Yang Zhang,Yougen Yuan,Yuming Yan,Jinjiang Liu,Zhiyong Wu,Pengfei Hu,Qun Yu.Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding[EB/OL].(2025-05-21)[2025-06-21].https://arxiv.org/abs/2505.15380.点此复制

Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding

Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding

评论