|国家预印本平台
首页|Probing the Robustness Properties of Neural Speech Codecs

Probing the Robustness Properties of Neural Speech Codecs

Probing the Robustness Properties of Neural Speech Codecs

来源:Arxiv_logoArxiv
英文摘要

Neural speech codecs have revolutionized speech coding, achieving higher compression while preserving audio fidelity. Beyond compression, they have emerged as tokenization strategies, enabling language modeling on speech and driving paradigm shifts across various speech processing tasks. Despite these advancements, their robustness in noisy environments remains underexplored, raising concerns about their generalization to real-world scenarios. In this work, we systematically evaluate neural speech codecs under various noise conditions, revealing non-trivial differences in their robustness. We further examine their linearity properties, uncovering non-linear distortions which partly explain observed variations in robustness. Lastly, we analyze their frequency response to identify factors affecting audio fidelity. Our findings provide critical insights into codec behavior and future codec design, as well as emphasizing the importance of noise robustness for their real-world integration.

Wei-Cheng Tseng、David Harwath

通信无线通信电子技术应用

Wei-Cheng Tseng,David Harwath.Probing the Robustness Properties of Neural Speech Codecs[EB/OL].(2025-05-30)[2025-06-25].https://arxiv.org/abs/2505.24248.点此复制

评论