arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08210cs.AI

对齐的错觉:检测协同对话中的隐藏分歧

Illusion of Alignment: Detecting Hidden Disagreement in Collaborative Dialogue

Kaiming Liu, Fuwen Luo, Ziyue Wang, Jinrui Ju, Yuxuan Liu, Xuanyu Lei, Yunghwei Lai, Peng Li, Yang Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对协同对话中存在的对齐的错觉(IoA)问题,构建IoA-Suite数据集,训练IoA-Prober-8B模型检测隐藏分歧,该模型可揭示真人会议中未表达的分歧,还能提升多智能体协作的下游任务性能。

中文摘要 AI 辅助

协同对话可能以表面一致告终,但参与者在目标、假设或执行计划上仍存在差异,从而形成“对齐的错觉(IoA)”。一项涵盖18场会议的真人研究证实,IoA在人类协作中频繁出现。然而IoA构成了一个悖论:若参与者意识到此类分歧,它们本就已明确表达;若未意识到,被询问时也无法阐明,致使IoA对参与者和观察者均不可见。本研究通过生成诊断性多项选择题,使IoA可被检测,参与者对这些问题的不同回答为隐藏分歧提供了直接行为证据。我们构建了IoA-Suite——一个用于检测隐藏分歧的数据集与评估协议,涵盖5种任务类型和6个领域。研究发现,即便最优模型也仅达到49.5%的F1值,瓶颈源于对话未呈现的私人上下文。随后我们基于IoA-Suite训练了IoA-Prober-8B,在IoA-Suite上达到51.8%的F1值。在上述18场真人会议中,它每场会议可揭示2.89个参与者确认未表达的隐藏分歧,且可迁移至实时人类对话。此外,在多智能体协作中,将IoA-Prober-8B与大语言模型智能体配对,可提升BigCodeBench-Hard和HiddenBench上的下游任务性能。

英文摘要

Collaborative dialogue can end with apparent agreement while participants still differ on goals, assumptions, or execution plans, creating an \textbf{illusion of alignment (IoA)}. A real-user study across 18 meetings confirms that IoA arises routinely in human collaboration. Yet IoA poses a paradox: if participants were aware of such disagreements, they would already be explicit; if not, they cannot articulate them when asked, leaving IoA invisible to both participants and observers. In this work, we make IoA detectable by generating diagnostic multiple-choice questions whose divergent answers across participants provide direct behavioral evidence of hidden disagreement. We construct \textbf{IoA-Suite}, a dataset and evaluation protocol for detecting hidden disagreement, spanning five task types and six domains. We find that even the best model attains only 49.5\% F1, with the bottleneck traced to private context that the dialogue does not surface. We then train \textbf{IoA-Prober-8B} based on IoA-Suite, reaching 51.8\% F1 on IoA-Suite. Across the aforementioned 18 real meetings, it surfaces 2.89 hidden disagreements per meeting that participants confirm they had not voiced, transferring to live human dialogue. Further, in multi-agent collaboration, pairing IoA-Prober-8B with LLM agents improves downstream task performance on BigCodeBench-Hard and HiddenBench.

发表机构

  • College of AI, Tsinghua University(清华大学人工智能学院)
  • Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)
  • Institute for AI, Tsinghua University(清华大学人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑