发表机构
Centific Research; University of Washington(Centific研究院; 华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
通过受控实验解耦声音与内容,发现S2S模型在性别归因上受内容刻板印象主导,而非声音,且固定声音评估无法察觉此偏见。
AI 中文摘要
语音到语音(S2S)模型现已应用于配音、翻译和语音代理中。与文本模型不同,它们能听到说话者的声音,而声音承载着说话者的性别。一个忠实的系统应将说话者视为其听起来的样子,而非通常说出其话语内容的人。测试这一点比看起来更难,因为大多数S2S模型以单一、固定的输出声音作答,该声音被硬编码,不会向刻板印象漂移。即使模型存在偏见,检查输出声音也会得到干净的结果。因此,我们提出两个问题:当模型重新说出输入内容时,词语中的刻板印象是否会改变输出声音的感知性别(声音渲染)?当模型陈述说话者性别时,它是依据声音还是内容(性别归因)?我们通过一个受控实验回答这两个问题,该实验将男性和女性声音与男性化、中性和女性化刻板印象的段落交叉组合,在英语、西班牙语和普通话的五种开源和闭源模型上进行。渲染后的声音未显示刻板印象漂移。但每个模型都根据内容而非声音来决定说话者的性别。将内容向女性化方向提升一个级别(男性化→中性→女性化)会使“女性”判断的几率乘以1.7至24倍。当内容与声音冲突时,最差的模型在90%的情况下错误判断说话者性别;当它们一致时,错误率仅为2%。因此,偏见隐藏于性别归因中,而固定声音评估无法察觉,随着S2S系统越来越多地为真实人群发声,审计必须关注这一环节。
英文摘要
Speech-to-speech (S2S) models now run inside dubbing, translation, and voice agents. Unlike text models, they hear the speaker's voice, which carries the speaker's gender. A faithful system should treat a speaker as who they sound like, not as whoever usually says what they said. Testing this is harder than it looks, since most S2S models answer in a single, fixed output voice, hard-coded so it cannot drift toward a stereotype. Checking the output voice comes back clean even when the model is biased. We therefore ask two questions. When a model re-speaks the input, does the stereotype in the words shift the perceived gender of the output voice (voice rendering)? And when the model states the speaker's gender, does it follow the voice or the content (gender attribution)? We answer both with one controlled experiment crossing male and female voices with masculine-, neutral-, and feminine-stereotyped passages, on five open- and closed-source models in English, Spanish, and Mandarin. The rendered voice shows no stereotype drift. But every model decides the speaker's gender from the content, not the voice. Making the content one step more feminine (masculine -> neutral -> feminine) multiplies the odds of a "female" judgment by 1.7-24. When the content clashes with the voice, the worst model misgenders the speaker in 90% of cases. When they agree, it misgenders in only 2%. The bias thus hides in gender attribution, where fixed-voice evaluation cannot see, and where audits must look as S2S systems increasingly speak for real people.