arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38604cs.CLcs.AIcs.LG

超越完美通信:在误解与用户意图演变下的交互式意图对齐基准测试

Beyond Oracle Communication: Benchmarking Interactive Intent Alignment Under Miscommunication and Evolving User Intent

  • ByteDance Inc.(字节跳动公司)
  • University of Notre Dame(圣母大学)

机构由 AI 辅助整理,请以论文原文为准。

Zheyuan Zhang, Mengyuan Chao, Ke Xiao, Ziyi Chen, Daoan Zhang, Yan Zhang, Yanfang Ye, Wei Xu

AI总结:

针对现有基准假设用户完美通信的局限,提出Drift-Bench++基准和GRIP评估协议,研究智能体在不完美通信和意图演变下的交互式意图对齐,发现更强交互有帮助但远未达到完美性能。

AI中文摘要:

现代LLM智能体越来越多地通过与用户的交互式、长时程交流来处理复杂任务,而现有基准通常假设用户总是能够准确且充分地传达一个固定的意图。然而,这种完美通信假设在实践中很少成立:用户可能会误解、改变目标,并失去耐心。我们将此任务设定定义为交互式意图对齐,即智能体必须在不完美通信和不断演变的目标下,恢复并持续追踪用户的当前意图。为了研究这一设定,我们引入了Drift-Bench++,一个用于生成具有受控误解和意图转移的可验证可执行任务的原则性基准构建流程,以及一个包含有限耐心、多样化模拟用户和静默交互条件转移的交互协议。我们进一步开发了GRIP,一个涵盖任务基础、用户真实性、询问有效性和对演变意图适应性的综合评估协议。在多样化的环境、模型和交互条件下,更强的交互始终有所帮助,但远未达到完美通信的性能;在部署的ProdAgent会话上的验证进一步表明,所建模的失败在部署中普遍存在且影响重大。通过为交互式意图对齐提供一个统一、可执行的基准,Drift-Bench++为在现实通信和演变意图下评估和推进智能体奠定了基础。

英文摘要:

Modern LLM agents increasingly tackle complex tasks through interactive, long-horizon exchanges with users, while existing benchmarks generally assume that users always accurately and sufficiently communicate a fixed intent. However, this oracle communication assumption rarely holds in practice: users may miscommunicate, change their goals, and run out of patience. We define this task setting as Interactive Intent Alignment, where agents must recover and continuously track the user's current intent despite imperfect communication and evolving goals. To study this setting, we introduce Drift-Bench++, a principled benchmark construction pipeline for verified executable tasks with controlled misalignment and intent shifts, along with an interaction protocol featuring finite patience, diverse simulated users, and silent interaction-conditioned shifts. We further develop GRIP, a comprehensive evaluation protocol covering task grounding, user realism, inquiry effectiveness, and adaptation to evolving intent. Across diverse environments, models, and interaction conditions, stronger interaction consistently helps but remains far from oracle performance; Validation on deployed ProdAgent sessions further shows that the modeled failures are prevalent and consequential in deployment. By providing a unified, executable benchmark for interactive intent alignment, Drift-Bench++ offers a foundation for evaluating and advancing agents under realistic communication and evolving intent.

↑