arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HEAR:解锁说话人归因推理的反事实语音接地方法

HEAR Who Said What: Unlocking Speaker-Attributed Reasoning via Counterfactual Voice Grounding

Dongwook Lee, Sangkwon Park, Eunwoo Song, Che Hyun Lee, Youngho Cho, Junho Kim, June Young Yi, Heeseung Kim, Sungroh Yoon

arXiv 2608.29120首次发表:更新:

发表机构

IPAI, Seoul National University (SNU); Yonsei University; University of Seoul(首尔国立大学IPAI; 延世大学; 首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出说话人归因推理基准HEAR,发现现有SLMs依赖语义先验而非语音线索,进而提出30B模型A2R,在CASH数据集优化后于HEAR表现优异,且可零样本泛化至多说话人下游任务。

AI 中文摘要

语音语言模型(SLMs)越来越多地被部署在多说话人环境中,但它们将语音归因于正确说话人并基于说话人身份进行推理的能力仍不明确。因此,我们推出HEAR,这是一个概念分层的基准,用于诊断说话人归因推理的基础能力,包含来自887个不同多方音频片段的2400个人工验证样本。对20个领先的SLMs在HEAR上的评估显示,它们在这些基础任务上表现不佳,通常依赖语义先验而非实际语音线索。为解决此问题,我们提出A2R,这是一个30B参数的模型,在反事实音频与说话人级难负样本(CASH)上进行优化,CASH是一个旨在引导模型优先考虑声学语音线索而非语言信号的数据集。A2R在HEAR上取得了出色性能,并在不同的多说话人下游任务上表现出零样本泛化能力,表明学习到的说话人归因能解锁模型潜在的说话人感知推理能力。所有资源均可在指定URL获取。

英文摘要

Speech Language Models (SLMs) are increasingly deployed in multi-speaker environments, yet their ability to attribute speech to the correct speaker and reason over speaker identities remains unclear. Hence, we introduce HEAR, a conceptually hierarchical benchmark diagnosing the foundational capabilities of speaker-attributed reasoning, comprising 2.4K human-verified samples from 887 diverse multi-party audio clips. Evaluating 20 leading SLMs on HEAR reveals they struggle with these foundational tasks, often relying on semantic priors rather than actual vocal cues. To address this, we present A2R, a 30B model optimized on Counterfactual Audio with Speaker-level Hard negatives (CASH), a dataset designed to guide the model to prioritize acoustic vocal cues over linguistic signals. A2R achieves strong performance on HEAR and exhibits zero-shot generalization to diverse multi-speaker downstream tasks, demonstrating that learned speaker attribution unlocks the model's latent capacity for speaker-aware reasoning. All resources are available at https://attributetoreason.github.io/AttributeToReason/

CommentsEMNLP2026 Main Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑