音频语言模型是否使用副语言证据?用于响应评估的反事实审计
Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation
浏览论文内容
中文总结 AI 辅助
该研究针对用作语音转语音系统评判者的音频语言模型,提出反事实审计方法,发现其评判可靠性常被对比成功高估,需结合行为审计而非仅准确率评估。
中文摘要 AI 辅助
音频语言模型(ALM)越来越多地被用作语音转语音系统的评判者,但接收音频的评判者可能实际上并未使用副语言证据。我们引入用于副语言响应评估的反事实审计:每个审计项固定文本,同时改变情感、韵律或情感转变的时间,迫使有效评判者追踪音频线索而非词汇内容或响应风格。我们使用原生单上下文判断协议和对比可恢复性控制评估ALM评判者,再将每个项分解为其组成的感知和响应映射技能,这产生了可识别不同评判失败来源的有用诊断状态。在Gemini、GPT和开源音频模型中,我们发现对比成功常常高估原生评判可靠性,且相似的总准确率可能隐藏不同的失败模式。这些结果表明,不应仅通过准确率评估ALM评判者,部署前需进行彻底的行为审计。
英文摘要
Audio-language models (ALMs) are increasingly used as judges for speech-to-speech systems, but a judge that receives audio may not actually use paralinguistic evidence. We introduce counterfactual audits for paralinguistic response evaluation. Each audit item holds the transcript fixed while varying affect, prosody, or the timing of an affective shift, forcing a valid judge to track the audio cue rather than lexical content or response style. We evaluate ALM judges using a native one-context judgment protocol and a contrastive recoverability control, then further decompose each item into its constituent perception and response-mapping skills. This yields useful diagnostic states that identify different sources of judge failures. Across Gemini, GPT, and open audio models, we find that contrastive success often overstates native judge reliability, and that similar aggregate accuracies can hide different failure modes. These results suggest that ALM judges should not be evaluated by accuracy alone, instead requiring thorough behavioral audits before deployment.
发表机构
- Boston University(波士顿大学)
机构由 AI 辅助整理,请以论文原文为准。