AI 中文总结
研究大型音频语言模型语音评估中协议级捷径,通过对三种部署协议审计发现多个LALM依赖此类捷径,如特征蓝图评判中错误标签影响准确率,拼接A/B比较有固定选择倾向,强调联合评估模型与协议及用匹配探针评估的重要性。
AI 中文摘要
大型音频语言模型(LALMs)越来越多地被用作语音评估的自动评判工具。然而,与人类评分的高度一致性并不能保证其评判基于音频。评判者可能会依赖评估协议提供的专业标签或参考数据,而不是听音频。本文在三种常见部署协议中审计LALM评判中的协议级“捷径”:特征蓝图评判(音频被声学特征的结构化文本描述取代)、参考条件评判和成对A/B比较。在六个评判者和四个属性上,发现多个LALM依赖协议级捷径。如在特征蓝图评判中,错误的专业标签使五个评判者的情感准确率降至0.10或更低;在拼接A/B比较中,Qwen3 - Omni - Thinking常选相同槽位。结果表明,除非联合评估模型和评估协议,否则总体一致性可能高估LALM评判的有效性,且每个模型 - 协议对都应用匹配的捷径探针评估。
英文摘要
Large audio-language models (LALMs) are increasingly used as automatic judges for speech evaluation. However, high agreement with human ratings does not guarantee that their verdicts are grounded in the audio. A judge may instead rely on specialist labels or reference data supplied by the evaluation protocol itself, taking a shortcut in place of listening to the audio. In this paper, we audit such protocol-level ``shortcuts'' in LALM judges across three common deployment protocols: feature-blueprint judging, where the audio is replaced by a structured text description of acoustic features, reference-conditioned judging, and pairwise A/B comparison. Across six judges and four attributes, we find that several LALMs rely on protocol-level shortcuts. For example, in feature-blueprint judging, incorrect specialist labels reduce five judges' emotion accuracy to 0.10 or below, and in concatenated A/B comparisons, Qwen3-Omni-Thinking often picks the same slot regardless of order swaps. These results indicate that aggregate agreement can overstate the validity of LALM judges unless the model and the evaluation protocol are assessed jointly, and that each model-protocol pair should be evaluated with a matched shortcut probe.