带着地图迷路:语言模型中的对话状态与行为可靠性
Lost with a Map: Conversational State and Behavioral Reliability in Language Models
浏览论文内容
中文总结 AI 辅助
本研究探究语言模型在任务导向对话中的状态表示与更新机制,发现结构与数值分离,并提出一种基于结构读数的状态-动作控制器,显著提升查询准确率与任务成功率。
中文摘要 AI 辅助
任务导向型对话需要在多轮对话中维护和更新信息,然而语言模型并未暴露任何明确的信念状态对象。我们研究了在来自四个家族的八个指令微调语言模型中,对话状态如何在MultiWOZ和SGD数据集上被表示、更新和使用。结构与数值是分离的:哪些领域、槽位和请求处于活跃状态,在模型即将行动之前是线性可读的,而精确数值在用户陈述它们的位置则更易读。当用户更改某个数值后,两个数值在其提及位置仍然可访问,因果干预表明两者继续影响模型的动作。在自然的闭环交互中,查询失败可分为结构性支持薄弱、数值解析错误以及未能部署本应受支持的约束等情况,针对性的干预在这些情况中产生了系统性不同的修复行为。这些发现催生了一种状态-动作控制器,它从基础模型的动作出发,利用结构读数进行选择性编辑,而无需将完整的预测信念状态作为中间表示。在五个模型的保留MultiWOZ交互中,它将基础模型的精确查询准确率从0.318提升到0.621,任务成功率从0.272提升到0.371,且附加成本可忽略不计。总体而言,可靠的交互不仅需要保留对话信息,还需要解析当前哪些可用约束适用,并确保它们支配动作。
英文摘要
Task-oriented dialogue requires maintaining and updating information across turns, yet language models expose no explicit belief-state object. We study how conversational state is represented, updated, and used inside eight instruction-tuned language models from four families on MultiWOZ and SGD. Structure and values separate: which domains, slots, and requests are active is linearly readable just before the model acts, whereas exact values are far more readable where the user stated them. After a user changes a value, both values remain accessible at their mentions, and causal interventions show that both continue to influence the model's action. In natural closed-loop interaction, query failures separate into cases of weak structural support, incorrect value resolution, and failure to deploy otherwise-supported constraints, with targeted interventions producing systematically different repair behavior across these cases. These findings motivate a state-action controller that starts from the base model action and selectively edits it using structural readouts, without requiring a complete predicted belief state as an intermediate representation. On held-out MultiWOZ interaction across five models, it raises the base model exact-query accuracy from .318 to .621 and task success from .272 to .371 at negligible added cost. Overall, reliable interaction requires not only retaining conversational information, but resolving which available constraints currently apply and ensuring that they govern action.
发表机构
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。