发表机构
Center for Language and Speech Processing, Johns Hopkins University(约翰霍普金斯大学语言与语音处理中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对口语对话理解中基于转录本的捷径问题,构建了含501个问题的受控基准ContraTalk,提出Audio Twin智能体式推理框架,可提升冲突问答案例的准确率并减少文本偏向陷阱选择。
AI 中文摘要
理解口语对话需要对词汇内容和副语言声学信号(如情感和对话意图)进行联合推理。然而,现有评估常允许基于文本转录或单模态解决方案的捷径,模糊了模型是否真正将预测建立在语音基础上。我们将这种失败模式形式化为跨模态不一致,即转录本提供看似合理但错误的表面解释,而韵律或说话风格等声学线索支持不同答案。我们开发了一个可扩展框架,用于识别文本偏向的表面解释,并将不一致区域转换为冲突问答示例。我们还纳入了转录本与语音接地解释一致的案例,实现了超越对抗性音频依赖的评估。由此得到ContraTalk,这是一个受控基准,包含501个问题,涵盖五个话语维度:交互行为、情感状态、对话行为、社会立场和对话意图。我们进一步开发了一种智能体式推理框架,将语音转换为Audio Twin,即本地化声学线索的文本可读表示,向推理模型暴露声学证据。实验表明,强大的仅文本大型语言模型(LLM)在一致案例中准确率超过90%,但在冲突案例中降至33%-48%;直接AudioLLM仅提供部分接地,仍在约30%-40%的冲突案例中选择文本偏向陷阱;我们的Audio Twin框架提高了冲突案例的准确率,同时减少了陷阱选择,但其在一致案例中的表现仍依赖于 backbone。这些结果表明,基于转录本的捷径是口语对话理解中的重要失败模式,而显式声学证据聚合为诊断和改进语音接地推理提供了更可控的接口。
英文摘要
Understanding spoken dialogue requires joint reasoning over lexical content and paralinguistic acoustic signals such as emotion and conversational intent. However, existing evaluations often allow shortcuts based on transcripts or single-modality solutions, obscuring whether models genuinely ground predictions in speech. We formalize this failure mode as cross-modal disagreement, where transcripts suggest plausible but incorrect surface interpretations while acoustic cues such as prosody or speaking style support different answers. We develop a scalable framework that identifies text-biased surface interpretations and converts disagreement regions into conflict QA examples. We also include consistent cases where transcript-based and speech-grounded interpretations agree, enabling evaluation beyond adversarial audio dependence. This results in ContraTalk, a controlled benchmark containing 501 questions across five discourse dimensions: interaction behavior, emotion state, dialogue act, social stance, and conversational intent. We further develop an agentic-style reasoning framework that converts speech into an Audio Twin, a text-readable representation of localized acoustic cues that exposes acoustic evidence to the reasoning model. Experiments show that strong text-only LLMs exceed 90% accuracy in consistent cases but drop to 33-48% in conflict cases. Direct AudioLLMs provide only partial grounding, still selecting the transcript-biased trap in roughly 30-40% of conflict cases. Our Audio Twin framework improves conflict-case accuracy while reducing trap selection, but its consistent-case behavior remains backbone-dependent. These results identify transcript-based shortcuts as an important failure mode in spoken dialogue understanding and show that explicit acoustic evidence aggregation provides a more controllable interface for diagnosing and improving speech-grounded reasoning.
Comments24 pages, 4 figures