arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36979eess.AScs.SD

更响、更长、更生动:语音LLM裁判中的声学捷径与欠规格理由

Louder, Longer, Livelier: Acoustic Shortcuts and Underspecified Rationales in Speech LLM Judges

发表机构伊利诺伊大学厄巴纳-香槟分校 · 网飞
查看机构详情
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • Netflix(网飞)

机构由 AI 辅助整理,请以论文原文为准。

Mingyue Huo, Shivam Mehta, Bhavin Jawade, Yinghong Lan, Haoqi Li

首次发表
浏览论文内容

中文总结 AI 辅助

本研究审计六个语音LLM裁判,发现其存在声学捷径(偏好更响、更丰富、更生动的语音),且理由缺乏声学依据,并提出可靠裁判需抵制捷径并基于声学证据给出理由。

中文摘要 AI 辅助

LLM作为裁判被广泛用于评估文本,但将这一范式扩展到语音需要模型同时解释声学证据和语言证据。这引入了一种模态特有的风险:语音裁判可能将感知上显著的线索视为质量的证据,即使该线索与目标标准无关,或者其权重高于人类听众所给予的。我们将这种行为称为声学捷径。为了研究它,我们使用对强度、内容丰富度和情感表达的可控操纵,审计了六个语音LLM裁判。我们评估了点式评分和成对比较,并使用人类偏好校准来解释结果。裁判一致地奖励更响亮的音频,比人类听众更强烈地偏好内容丰富的语音,并将情感表达映射到质量偏好中。这些效应在成对比较中最明显,而点式评分往往掩盖它们。更令人担忧的是,伴随的理由很少识别改变判断的声学线索,而是反复依赖有限的词汇,使它们在声学上欠规格。总之,这些发现表明,可靠的语音裁判必须既抵制声学捷径,又将其理由建立在决策背后的声学证据之上。为了支持可重复性和未来的审计,我们还发布了SpeechJudgeAudit,即本研究中使用的受控刺激和评估工具。

英文摘要

LLM-as-a-judge is widely used for evaluating text, but extending this paradigm to speech requires models to interpret acoustic as well as linguistic evidence. This introduces a modality-specific risk: a speech judge may treat a perceptually salient cue as evidence of quality even when that cue is irrelevant to the target criterion or receives more weight than human listeners give it. We call this behavior an acoustic shortcut. To study it, we audit six speech LLM judges using controlled manipulations of intensity, content richness, and emotional delivery. We evaluate both pointwise scoring and pairwise comparison, using human preference calibration to interpret the results. The judges consistently reward louder audio, prefer content-rich speech more strongly than human listeners do, and map emotional delivery into quality preferences. These effects are most visible in pairwise comparison, while pointwise scores often obscure them. More concerningly, the accompanying rationales rarely identify the acoustic cue that changes a judgment and instead repeatedly rely on a limited vocabulary, leaving them acoustically underspecified. Together, these findings show that reliable speech judges must both resist acoustic shortcuts and ground their rationales in the acoustic evidence behind their decisions. To support reproducibility and future audits, we also release SpeechJudgeAudit, the controlled stimuli and evaluation tools used in this study.

↑