arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10606cs.CL

ASR往返评估可掩盖中文新闻TTS中依赖上下文和惯例的朗读错误

ASR-Roundtrip Evaluation Can Mask Context- and Convention-Dependent Reading Errors in Chinese News TTS

Shijun Luo, Lizhi Wan

AI总结:

研究发现ASR往返评估可掩盖中文新闻TTS的依赖上下文和惯例的朗读错误,经MiMo、CosyVoice等审计及Qwen3-ASR、Paraformer验证,该评估仅适用于筛选,无法作为独立基准。

AI中文摘要:

ASR往返评估被广泛用作文本转语音(TTS)可懂度的可扩展替代指标,但它可能对听众感知到的朗读错误产生假阴性结果。本研究聚焦于中文新闻TTS中依赖上下文或领域惯例的朗读片段,如体育比分、飞机型号、技术单位和会员名称。在这些场景下,Raw TTS可能选择看似合理但错误的朗读方式,而ASR会将音频转录为预期或表面正确的文本。对110个高风险MiMo TTS案例进行的针对性审计(报告了完整分母)确认了46个被掩盖的假阴性、9个暴露的TTS错误以及55个无Raw TTS错误的案例。片段隔离诊断重新暴露了46个先前被掩盖错误中的18个。在同一针对性池上进行的仅Raw的CosyVoice审计确认了51个被掩盖的案例。在两次审计中标记为已确认被掩盖的97个TTS特定音频文件中,Qwen3-ASR表面恢复了40个案例,而Paraformer仅恢复了2个。结果表明,ASR往返评估可用于筛选,但不足以作为中文新闻朗读风险评估的独立基准。

英文摘要:

ASR-roundtrip evaluation is widely used as a scalable proxy for text-to-speech (TTS) intelligibility, but it can produce false negatives for reading errors perceived by listeners. We study Chinese news TTS spans whose correct reading depends on context or domain conventions, such as sports scores, aircraft models, technical units, and membership names. In these cases, Raw TTS can choose a plausible but wrong reading while ASR transcribes the audio as the intended or surface-correct text. A targeted audit over 110 high-risk MiMo TTS cases, reported with a complete denominator, confirms 46 masked false negatives, 9 exposed TTS errors, and 55 cases with no Raw TTS error. A span-isolation diagnostic re-exposes 18/46 previously masked errors. A Raw-only CosyVoice audit on the same targeted pool confirms 51 masked cases. Across the 97 TTS-specific audio files labeled confirmed masked across the two audits, Qwen3-ASR surface-recovers 40 cases, whereas Paraformer does so in only 2. The results suggest that ASR-roundtrip is useful for screening but insufficient as standalone ground truth for Chinese news reading-risk evaluation.

补充信息

↑