发表机构
Meta Superintelligence Labs(Meta超级智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
我们提出HEAR基准,包含87k个来自843名参与者的真实语音样本,用于评估音频大语言模型中的语音条件偏见,发现偏见是模型特定且可控的,个性化指令会加剧人口差异。
AI 中文摘要
我们引入了HEAR(由真实说话者进行的音频大语言模型偏见人类录音评估),这是一个大规模、生态有效的基准,包含来自843名人口多样性参与者的87k个真实人类音频样本。HEAR通过多项选择问答(MCQA)和开放式长形式任务实现全面评估。据我们所知,这是首个完全基于真实人类语音的大规模语音基准。我们评估了实时语音到语音和语音到文本两种架构下的模型行为。我们的结果显示,语音条件偏见是一种模型特定属性。此外,我们证明个性化指令持续加剧人口统计学差异。我们的发现确立了语音偏见是一种可控的模型特性,为未来音频大语言模型开发中的偏见缓解和评估提供了基础框架。
英文摘要
We introduce HEAR (Human-recorded Evaluation of Audio-LLM bias by Real speakers), a large-scale, ecologically valid benchmark comprising 87k real human audio samples from 843 demographically diverse participants. HEAR enables comprehensive evaluation through Multiple Choice Question Answering (MCQA) and open-ended long-form tasks. To our knowledge, this is the first large-scale voice benchmark grounded entirely in authentic human speech. We evaluate model behavior across both real-time speech-to-speech and speech-to-text architectures. Our results reveal that voice-conditioned bias is a model-specific property. Furthermore, we demonstrate that personalization instructions consistently exacerbate demographic disparities. Our findings establish that voice bias is a controllable model characteristic, providing a foundational framework for future bias mitigation and evaluation in Audio-LLM development.