针对ASR模型中基准优化的量化研究
Towards Quantifying Benchmark Optimization in ASR Models
浏览论文内容
中文总结 AI 辅助
本文提出量化ASR模型基准优化的方法,识别三类行为探针,发现高分开源模型存在基准优化行为,该行为可被操控且会虚增基准性能。
中文摘要 AI 辅助
公开基准是衡量自动语音识别(ASR)模型能力的重要指标,但由于其公开性,存在模型针对这些基准进行优化,导致无法良好泛化到真实世界数据的风险。本文提出一种量化基准优化的方法,重点关注音频无法确定参考转录的情况。我们识别出三类行为探针,可揭示模型在音频不确定时重现基准参考片段的能力:参考分歧、掩码数字恢复和正字法切换。研究发现,得分最高的开源ASR模型会输出逐字的参考转录片段,即使相关音频存在矛盾、被掩码或模糊不清。通过多种机制探针,我们表明模型会响应狭窄的声学线索,覆盖对音频的忠实表示,转而采用基准优化策略。研究还显示,基准优化行为可通过低秩线性操控,或在某些情况下仅在片段末尾追加音频来实现因果操控。总体而言,研究结果表明,高性能模型表现出依赖基准的行为,这类行为会虚增基准性能,却无法反映通用转录能力的提升。
英文摘要
Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature of being public, there is risk of models being optimized for these benchmarks in ways that do not generalize well to real-world data. We present a methodology for quantifying benchmark optimization, focusing on cases where the audio underdetermines the reference transcript. We identify three families of behavioral probes that reveal models' capabilities of reproducing benchmark reference spans despite underdetermined audio: reference disagreement, masked-number recovery, and orthographic switching. We find that the highest-scoring open source models output verbatim reference transcript spans even when the relevant audio is contradictory, masked, or ambiguous. Using a variety of mechanistic probes, we show that models respond to narrow acoustic cues to override the faithful representation of the audio in favor of a benchmark-optimized policy. We show the benchmark-optimized behavior can be causally manipulated via low-rank linear steering or simply appending audio to the end of a segment in some cases. Overall, our results indicate that high-performing models exhibit benchmark-conditioned behaviors that can inflate benchmark performance without reflecting improved general-purpose transcription ability.
发表机构
- Hume AI Research(休姆人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。