XYEval:智能体对糟糕建议说“是”
XYEval: Agents say yes to bad advice
浏览论文内容
中文总结 AI 辅助
本文提出XYEval元评估框架,将现有基准转为XY问题评估,发现智能体在误导建议下性能大幅下降(最高46.7%),且缓解困难,需同时识别误导并清晰沟通。
中文摘要 AI 辅助
用户与AI智能体之间的有效沟通对于人机协作至关重要。XY问题是一种众所周知的沟通陷阱,即人们询问的是他们尝试的解决方案而非实际问题。我们将先前的谄媚评估扩展到智能体环境中的XY问题,评估智能体是否能抵抗用户提出的看似合理但具有误导性的建议,并传达其推理过程。我们引入了XYEval,一个元评估框架,可以将现有基准转换为XY问题评估。我们评估了六个不同基准套件中的五个模型。智能体在XY变异下,跨基准遭受严重的XY性能下降,相对下降最高达46.7%。通过τ²-bench,我们进一步表明,当遇到要求详细解释后才批准更好解决方案的迂腐用户时,智能体性能下降更多。我们的发现表明,当前智能体在面对误导性建议时,缺乏有效推理和沟通的能力。一个简单的系统指令基线,鼓励对XY问题的意识,仅提供部分缓解。广泛的轨迹分析提供了关于这些XY下降在执行轨迹中如何及为何发生的行为洞察。我们的结果表明,缓解XY问题仍然具有挑战性,要求智能体既能识别用户的误导,又能清晰沟通潜在问题。
英文摘要
Effective communication between users and AI agents is essential for human-AI collaboration. The XY problem is a well-known communication pitfall where a person asks about their attempted solution rather than their actual problem. We extend prior sycophancy evaluation to the XY problem in agentic settings, evaluating whether agents can resist plausible but misleading suggestions from users and communicate their reasoning. We introduce XYEval, a meta-evaluation framework that can transform an existing benchmark into an XY problem evaluation. We evaluate five models across six diverse benchmark suites. Agents suffer large XY drops under XY mutation across benchmarks, with relative drops reaching up to 46.7%. With $τ^2$-bench, we further show that agent performance drops more when encountering a pedantic user who requires detailed explanations before approving a better solution. Our findings suggest that current agents lack the ability to effectively reason and communicate when facing misleading suggestions. A simple system instruction baseline that encourages awareness of XY problems only offers partial mitigation. Extensive trace analyses provide behavioral insights into how and why these XY drops occur across execution trajectories. Our results show that mitigating the XY problem remains challenging, requiring agents to both recognize user misdirection and clearly communicate the underlying problem.
发表机构
- Google DeepMind(谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。