发表机构
Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究语言模型在面对模糊用户任务时的任务对齐问题,引入基于部分可观测马尔可夫决策过程的框架,通过用户研究验证,发现模型任务对齐困难,训练能改善但仍落后于人类,揭示其缺乏关键交互能力。
AI 中文摘要
当前语言模型基准主要评估在完全指定任务上的执行情况。然而,实际用户任务往往模糊不清,用户目标不完整、具探索性甚至不一致,需要助手先确定预期任务再执行。我们将此问题作为任务对齐来研究,引入一个将指定任务转换为未完全指定交互的通用框架,形式化为部分可观测马尔可夫决策过程,模型须从部分且不断演变的用户意图中推断潜在任务。通过用户研究事后验证用户模拟器。结果表明,虽任务指定后模型表现良好,但在任务对齐方面仍有困难,平均仅22 - 32%能恢复用户意图任务,人类在相同设置下可达48%。训练后能改善任务对齐,但模型在通过交互解决不确定性方面仍落后于人类。
英文摘要
Current benchmarks for language models primarily evaluate execution on fully specified tasks. However, real user tasks are often ambiguous. Users arrive with incomplete, exploratory, or even inconsistent goals, requiring the assistant to first determine the intended task before carrying it out. We study this problem as task alignment: the ability to align with a user on their intended task. We introduce a general framework for converting specified tasks into underspecified interactions, formalized as a POMDP in which the model must infer a latent task from partial and evolving user intent. We validate our user simulator post hoc with a human user study. Across shopping, coding, and professional work settings, we find that while models often perform well once the task is specified, models still struggle with task alignment: current models act prematurely, interact ineffectively, and fail to resolve ambiguous requests. Models on average recover the user's intended task only 22-32% of the time under ambiguity. In a human study in the same setting, humans reach 48%, outperforming all evaluated models. We show that post-training with supervised fine-tuning and reinforcement learning improves task alignment, but models still lag behind humans in resolving uncertainty through interaction. Together, our results suggest that current models still lack key interaction abilities required for reliable agency.