arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

辅助智能体通过信息价值推理进行理性澄清

Rational Clarification by Assistive Agents via Value-of-Information Reasoning

T. Duy Nguyen-Hien, Yee Whye Teh, Wee Sun Lee, Tan Zhi-Xuan

arXiv 2609.37588首次发表:更新:

发表机构

National University of Singapore; University of Oxford; Agency for Science, Technology and Research (A*STAR)(新加坡国立大学; 牛津大学; 新加坡科技研究局)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出REVOIR方法,通过信息价值推理决定是否澄清模糊请求,在两项任务中提升成功率并减少提问次数,优于现有方法。

AI 中文摘要

基于语言的辅助智能体的用户经常提出模糊的请求。作为回应,助手可以直接根据其对请求的解释采取行动——冒着与用户不一致的风险——或者提出澄清问题。哪种选择最安全、最有帮助?一种常见的方法是提出能最小化用户意图不确定性的问题,直到达到某个阈值。然而,这种方法忽略了不确定性减少对下游性能的影响、提问与立即行动的成本,以及用户可能在未被询问的情况下提供纠正的可能性。为了应对这些权衡,我们引入了基于信息价值推理的理性探究(REVOIR)。REVOIR通过在推理时对问题的信息价值进行推理来做出澄清决策,这捕捉了因收到的答案而带来的任务奖励的预期改进。在两个辅助任务——模糊问答(CondAmbigQA)和偏好对齐的家庭任务规划(ADAPT)中,我们展示了REVOIR在比基于提示、思维链、微调或信息增益的方法更少的问题下取得了更大的成功,在ADAPT上将偏好满意度提高了13-15%,相比微调的澄清策略,同时无需训练且提问次数减少了五倍。此外,当助手在行动后能收到廉价的用户纠正时,REVOIR自然地推断出提问并不总是高效的,展示了我们方法的适应性。相比之下,我们发现普通的推理智能体未能自适应地澄清用户请求,并且随着推理努力的增加,请求的澄清次数反而减少。

英文摘要

Users of language-based assistive agents often make ambiguous requests. In response, an assistant can either directly act on its interpretation of the request --- risking misalignment with the user --- or ask a clarifying question. Which option is the most safe and helpful? A common approach is to ask questions that minimize uncertainty about the user's intent until a threshold is reached. However, this neglects the impact of uncertainty reduction on downstream performance, the costs of asking versus acting immediately, and the possibility that users may provide corrections without being asked. To navigate these trade-offs, we introduce Rational Enquiry via Value-of-Information Reasoning (REVOIR). REVOIR makes clarification decisions via inference-time reasoning about the value-of-information of a question, which captures the expected improvement in task reward due to the answer received. In two assistive tasks --- ambiguous question answering (CondAmbigQA) and preference-aligned household task planning (ADAPT) --- we show that REVOIR achieves greater success with fewer questions than approaches based on prompting, chain-of-thought, fine-tuning, or information gain, improving preference satisfaction on ADAPT by 13-15% over a fine-tuned clarification policy while requiring no training and asking five times fewer questions. Furthermore, when the assistant can receive cheap user corrections after acting, REVOIR naturally infers that asking questions is not always efficient, demonstrating the adaptivity of our approach. In contrast, we find that vanilla reasoning agents fail to adaptively clarify user requests, and request fewer clarifications as reasoning effort increases.

Comments54 pages, 11 figures. Under review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑