arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

他们彼此是谁?用于说话者关系推断的多智能体推理

Who Are They to Each Other? Multi-Agent Reasoning for Speaker Relationship Inference

Yaohan Guan, Yen-Ju Lu, Yuzhe Wang, Junhyeok Lee, Jesus Villalba, Laureano Moro Velazquez, Thomas Thebaud, Najim Dehak

arXiv 2609.09628首次发表:更新:

发表机构

Johns Hopkins University(约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出无需训练的多智能体推理框架,通过辩论和竞争协议从对话中推断说话者关系,在多数情况下优于基线,但声学线索利用不足。

AI 中文摘要

从口语对话中推断说话者关系是迈向具有社会意识的语音理解的重要一步。然而,这项任务仍未得到充分探索,且有监督建模的训练和扩展成本高昂。同时,现有的推理时大语言模型方法在处理可能支持多种合理解释的微妙、分散和多模态关系线索方面提供的结构有限。为解决这些局限性,我们引入了一个无需训练的多智能体推理框架,通过大语言模型智能体之间的结构化交互来组织推理,使得关系判断可以在无需任务特定训练的情况下被提出、质疑和裁决。我们通过两种互补设计实例化该框架。我们提出多角色多智能体辩论,作为标准多智能体辩论在说话者关系推断中的任务特定改编,为智能体分配互补角色或基于社会理论的视角,而非单一的未分化观点。相比之下,我们引入多智能体竞争,一种基于竞争的协议,通过成对裁决比较智能体判断,淘汰较弱候选,并保留最可辩护的一个。我们在无缝交互数据集上跨不同模态设置评估这些方法,涵盖二分类和细粒度关系细节预测。结果表明,在大多数情况下,它们优于零样本和现有基于多智能体的基线。人工评估进一步表明,这项任务即使对人类也具有挑战性。在包含文本的设置中,大语言模型方法有时能超越人工标注者,但在音频设置中竞争力较弱。综合来看,这些发现表明,关系推断受益于智能体之间结构化的推理时交互,而当前模型尚未完全捕捉声学线索。

英文摘要

Inferring speaker relationships from spoken conversations is an important step towards socially aware speech understanding. However, this task remains underexplored, and supervised modeling is costly to train and scale. At the same time, existing inference-time LLM approaches provide limited structure for handling subtle, distributed, and multimodal relational cues that may support multiple plausible interpretations. To address these limitations, we introduce a training-free multi-agent reasoning framework that organizes inference through structured interaction among LLM agents, allowing relationship judgments to be proposed, challenged, and adjudicated without task-specific training. We instantiate this framework with two complementary designs. We propose Multi-Role Multi-Agent Debate as a task-specific adaptation of standard multi-agent debate for speaker relationship inference, assigning agents complementary roles or social-theory-grounded perspectives rather than a single undifferentiated viewpoint. In contrast, we introduce Multi-Agent Compete, a competition-based protocol that compares agent judgments through pairwise adjudication, eliminates weaker candidates, and retains the most defensible one. We evaluate these methods on the Seamless Interaction dataset across different modality settings, covering both binary classification and fine-grained relationship-detail prediction. Results suggest that they improve over zero-shot and existing multi-agent baselines in most cases. Human evaluation further suggests that this task is challenging even for people. LLM methods can sometimes outperform human annotators in text-included settings but are less competitive in the audio setting. Together, these findings suggest that relationship inference benefits from structured inference-time interaction among agents, while acoustic cues are not yet fully captured by current models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑