arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

你的语音质量指标可被攻击程度如何?一个修正后的协议、一个基准,以及修补能带来什么

How Hackable Is Your Speech Quality Metric? A Corrected Protocol, a Benchmark, and What Patching Buys

Ali Alavi, Donald S. Williamson

arXiv 2610.10899首次发表:更新:

发表机构

The Ohio State University(俄亥俄州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对语音质量预测器的可被攻击程度,提出修正协议与基准,发现不同预测器可被攻击程度差异大,且攻击-检测-修补循环的强化效果有限,跨域性能损失明显。

AI 中文摘要

语音质量预测器越来越多地被用作奖励,但目前尚无公认的衡量其可被攻击程度的指标。常规测量存在两个缺陷:其一,扰动通过处理链(此处为神经编解码器)传递至预测器,该处理链自身会改变分数,而针对原始输入的评分会将此计入攻击,若改为参考未受扰动的往返过程,测得的可被攻击程度变化可达4倍(某防御措施下从0.31变为0.08);其二,一个训练好的攻击者只是一个样本,而非测量值:仅随机种子不同的5个攻击者,对同一固定预测器的成功率从0.00到0.38不等,因此防御主张需采用多个攻击者的最坏情况。在该协议下,4个已发表的预测器差异显著:NISQA在90%的话语上被攻击,SSL-MOS为21%,DNSMOS为14%,UTMOS为6%。随后,作者对一个封闭的攻击-检测-修补循环进行审计:它仅在自身攻击空间中强化预测器,强化程度小于攻击者之间的差异;随机扰动基线与其表现相当;跨域时系统SRCC损失高达0.30;针对修补后的预测器进行后训练的增强器,攻击它们的难度显著降低(PESQ为-0.03,而未修补时为-0.23)。代码、预注册信息及运行输出均已公开。

英文摘要

Speech quality predictors are increasingly used as rewards, yet no agreed measure of their hackability exists. The usual measurement has two flaws. First, the perturbation reaches the predictor through a processing chain -- here a neural codec -- that shifts the score on its own, which scoring against the raw input charges to the attack. Referencing the unperturbed round trip instead changes measured hackability by up to a factor of four (0.31 to 0.08 for one defence). Second, one trained attacker is a sample, not a measurement: five attackers differing only in random seed reach success rates from 0.00 to 0.38 against one fixed predictor, so a defence claim needs the worst case over several. Under this protocol, four published predictors differ widely: NISQA is hacked on 90% of utterances, SSL-MOS on 21%, DNSMOS on 14% and UTMOS on 6%. We then audit a closed attack-detect-patch loop. It hardens the predictor only in its own attack space, by less than the spread between attackers; a random-perturbation baseline matches it; and it costs up to 0.30 system SRCC out of domain. Enhancers post-trained against patched predictors hack them far less (PESQ -0.03 versus -0.23). Code, preregistration and run outputs are released.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑