AI 中文总结
本研究提出IntentFlux基准和StateForge方法,测量并部分修复LLM智能体在多轮交互中的意图漂移问题,实验显示显式状态维护可将平均任务得分从0.367提升至0.467。
AI 中文摘要
LLM智能体通常在多轮交互中运行,用户意图可能在执行前发生变化。我们研究意图漂移:即用户意图中被取代的部分继续影响最终答案或工具操作的失败模式。我们引入了IntentFlux,一个可执行的基准,它将可验证的任务转换为带有受控意图变化的对话,同时保留其原始评分器。在627个案例的校准中,随着对话包含更多被取代和撤回的信息,平均任务得分从0.476降至0.384。在八个模型中,当相同的最终任务必须从不断演变的对话中恢复,而不是在单轮中直接给出时,完全正确解决方案的比率显著更低。我们进一步引入了StateForge,它在生成前显式维护活动需求。在General-Test上,它将平均任务得分从0.367提高到0.467。提供真实最终状态进一步提高了性能,但仍未恢复到单轮性能,这表明状态估计误差仅解释了部分差距。这些结果确立了意图漂移作为一种可测量的多轮失败模式,并表明显式状态维护是一种部分缓解措施。
英文摘要
LLM agents often operate over multi-turn interactions in which user intent changes before execution. We study intent drift: the failure mode in which superseded parts of the user's intent continue to influence the final answer or tool action. We introduce IntentFlux, an executable benchmark that converts verifiable tasks into dialogues with controlled intent changes while preserving their original graders. In a 627-case calibration, mean task score falls from 0.476 to 0.384 as dialogues contain more superseded and withdrawn information. Across eight models, the rate of fully correct solutions is significantly lower when the same final task must be recovered from an evolving dialogue rather than given directly in a single turn. We further introduce StateForge, which explicitly maintains the active requirements before generation. On General-Test, it improves mean task score from 0.367 to 0.467. Providing the ground-truth final state improves performance further but still does not recover single-turn performance, indicating that state-estimation errors explain only part of the gap. These results establish intent drift as a measurable multi-turn failure mode and explicit state maintenance as a partial mitigation.
Comments15 pages, 3 figures