AI 中文总结
研究大语言模型在用户意图演变时的表现,引入框架将静态任务转换为动态对话,发现模型在静态设置下的良好性能无法迁移到意图演变场景,揭示了当前大语言模型在跟踪和响应演变意图方面的不足。
AI 中文摘要
随着大语言模型能力增强,越来越多地被用作协作代理通过迭代交互承担用户委托任务。但真实交互是动态的,用户很少预先明确意图,而是在对话中披露、修改和重塑意图。尽管如此,大语言模型仍主要在单轮、完全指定的设置中评估或训练。为此引入一个框架,将静态单轮任务转换为动态多轮对话,用户意图在对话中演变,同时保留每个任务的原始评估协议。通过多个任务发现,强大的静态设置性能无法转移到意图演变设置,各模型家族性能大幅下降。研究结果表明当前大语言模型尚未忠实跟踪和响应用户不断演变的意图,这对未来协作代理至关重要但静态评估无法体现。
英文摘要
As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks through iterative interaction. Yet genuine interaction is inherently dynamic: users rarely specify their intent upfront, instead disclosing, revising, and reshaping it as the conversation unfolds. Despite this, LLMs are still predominantly evaluated or trained in single-turn, fully-specified settings, leaving open a fundamental question: how well do LLMs track and act on user intent as it evolves over the course of a conversation? To study this, we introduce a framework that transforms static, single-turn tasks into dynamic multi-turn conversations in which the user's intent evolves across turns--incrementally revealed, revised, and at times redirected mid-conversation--while preserving each task's original evaluation protocol, enabling existing benchmarks to be reused as controlled testbeds without new annotation. Across multiple tasks, we surface a consistent phenomenon: strong static-setting performance does not transfer to the evolving-intent setting, with substantial drops across model families. Our findings point to a fundamental gap: today's LLMs do not yet faithfully track and act on the user's evolving intent, a capability invisible to static evaluation yet critical for future collaborative agents.
Comments20 pages, 10 figures