发表机构
BITS Pilani; MIT; Dartmouth College(比拉理工学院; 麻省理工学院; 达特茅斯学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究将AI助手澄清不完整意图问题重构为信息价值问题,提出强化学习框架,在图像生成中通过多轮模拟用户优化交互,实验表明能显著提升用户匹配效果并减少交互成本。
AI 中文摘要
AI助手接收到的请求常常遗漏了达成良好结果所需的信息,例如关于用户偏好或目标的信息。在这种情况下,助手必须在继续之前进行推测或询问更多信息。我们将此重新概念化为一个信息价值问题:助手应获取那些缺失会导致用户效用最大可避免损失的信息。这种信息事先很少已知;相反,助手必须预测它,以便最优地分配有限的用户交互。我们在图像生成中实例化此问题,并推导出一个使用多轮模拟用户的强化学习框架,以在不确定性下最大化效用恢复。在一项预注册研究中,涉及76名人类参与者的456次交互会话,这帮助用户显著更好地匹配参考图像,同时显著减少问题数量、总交互时间和成本。这指向一个简单且可扩展的框架,用于训练语言模型助手通过提出更具信息性的问题来更好地消除用户意图的歧义。
英文摘要
AI assistants receive requests that leave out information needed for a good outcome, for example about users' preferences or goals. They must then either speculate or ask for more information before proceeding. We reconceptualize this as a value-of-information problem: the assistant should acquire information whose absence causes the greatest avoidable loss in user utility. This is rarely known ex ante; rather, assistants must predict it in order to optimally allocate limited user interactions. We instantiate this problem in image generation and derive a reinforcement learning framework using multi-turn simulated users to maximize utility recovery under uncertainty. In a preregistered study with 456 interactive sessions across 76 human participants, this helped users significantly better match reference images with significantly fewer questions, less total interaction time, and lower cost. This points toward a simple and scalable framework for training language model assistants to better disambiguate user intent by asking more informative questions.
Comments43 pages, 15 figures