arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

你改变了主意,模型却没有:解析多轮对话中的意图

You Changed Your Mind, The Model Didn't: Demystifying Intent in Multi-Turn Dialogue

Junle Chen, Wei Chen, Zhengjun Huang, Zhoujin Tian, Yuxuan Liu, Kai Wang, Rui Chen, Xiaofang Zhou

arXiv 2610.06496首次发表:更新:

发表机构

HKUST; Tencent(香港科技大学; 腾讯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出Intent-Eval基准揭示语言模型在用户意图变化时易受“提及即生效”混淆影响,并开发Intent-OPSD框架,通过决策条件的自蒸馏提升模型遵循最终意图的能力。

AI 中文摘要

当大型语言模型处理多轮任务时,用户提出更改但最终拒绝该更改,模型应继续执行任务,仿佛一切未变。我们发现了一个令人惊讶的失败:仅仅提及被拒绝的更改就可能使任务执行偏离轨道,即使用户的最终意图保持不变。为了系统研究语言模型在用户意图演变下的行为,我们引入了Intent-Eval,一个涵盖工具操作、代码、数据库和数学的受控基准。在多样化的任务中,模型对被拒绝的提议和被取代的需求均表现脆弱,这与“提及即生效”的混淆一致:对话内容在被拒绝或替换后仍被视为有效需求。随着交互的继续,准确率下降可能加深或持续,凸显了区分“已提及内容”与“仍有效内容”的必要性。基于这一见解,我们提出了Intent-OPSD,一种决策条件的在策略自蒸馏框架,其中教师和学生模型从同一模型初始化。冻结的教师模型根据匹配用户决策的完整任务提供主动意图监督,训练学生模型在整个对话中遵循反映用户意图的主动需求。

英文摘要

When a large language model handles a multi-turn task and a user proposes a change but ultimately rejects it, the model should continue as if nothing changed. We find a surprising failure: merely mentioning a rejected change can derail task execution, even when the user's final intent remains unchanged. To systematically study language model behavior under evolving user intent, we introduce Intent-Eval, a controlled benchmark spanning tool actions, code, databases, and mathematics. Across diverse tasks, models are vulnerable to both rejected proposals and superseded requirements, consistent with mentioned-as-in-effect confusion: conversational content is treated as active requirements even after it has been rejected or replaced. Accuracy degradation can deepen or persist as interaction continues, highlighting the need to distinguish what has been mentioned from what remains in effect. Building on this insight, we propose Intent-OPSD, a decision-conditioned on-policy self-distillation framework with Teacher and Student initialized from the same model. The frozen Teacher provides active-intent supervision from the complete task matching the user's decision, training the Student on the full dialogue to follow active requirements reflecting user intent.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑