arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自动语音识别系统的生成式测试

Generative Testing of Automated Speech Recognition Systems

Yanis Xabier Wilbrand Peña, Oliver Weißl, Andrea Stocco

arXiv 2607.09833首次发表:更新:

发表机构

Technical University of Munich; fortiss GmbH(慕尼黑技术大学; fortiss公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对自动语音识别系统易受对抗性操纵问题,提出黑盒测试方法GATAS,通过在文本转语音模型音素级潜在空间操作生成失败输入,经实验评估其成功率高、失真低、感知质量高,证明无目标潜在空间优化可有效生成测试用例。

AI 中文摘要

自动语音识别(ASR)系统借助基于Transformer的模型实现了高精度,得以应用于关键领域。然而,它们仍易受对抗性操纵影响,尤其在黑盒环境中,攻击须保持感知自然度。本文介绍了GATAS,一种黑盒测试方法,通过在文本转语音模型的音素级潜在空间中操作来生成导致失败的输入。该方法不直接扰动波形,而是对潜在表示进行插值以引发转录错误,同时保持在自然语音流形内。攻击被表述为平衡语义差异和感知质量的多目标优化问题。针对白盒和黑盒基线的实证评估表明,GATAS成功率达98%,失真更低,感知质量更高,人类研究也证实了这一点。尽管无梯度访问,GATAS与白盒方法相比仍具竞争力,凸显表示和感知对齐比访问模型内部更关键。总体而言,结果表明无目标的潜在空间优化能有效为ASR系统生成现实且有效的测试用例。

英文摘要

Automatic speech recognition (ASR) systems have achieved high accuracy with transformer-based models, enabling deployment in critical applications. However, they remain vulnerable to adversarial manipulation, particularly in black-box settings where attacks must preserve perceptual naturalness. This work introduces GATAS, a black-box testing approach that generates failure inducing inputs by operating in the phoneme-level latent space of a text- to-speech model. Instead of perturbing waveforms directly, the approach interpolates latent representations to induce transcription errors while remaining within the manifold of natural speech. The attack is formulated as a multi-objective optimization problem balancing semantic divergence and perceptual quality. Our empirical evaluation against both white-box and black-box baselines shows that GATAS achieves a 98% success rate while producing lower distortion and higher perceptual quality, as confirmed by human studies. Despite operating without gradient access, GATAS remains competitive against white-box methods, highlighting that representation and perceptual alignment are more critical than access to model internals. Overall, our results demonstrate that untargeted latent-space optimization enables the efficient generation of realistic and effective test cases for ASR systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑