澄清用户还是验证世界?面向主动式智能体的不确定性路由
Clarify the User or Verify the World? Uncertainty Routing for Proactive Agents
浏览论文内容
中文总结 AI 辅助
提出PROUR框架,将智能体不确定性路由到澄清用户或验证环境,在τ-bench上提升成功率4.57%并减少交互步骤,且能泛化到新领域。
中文摘要 AI 辅助
使用工具的LLM智能体不仅需要决定是否需要额外信息,还需要决定哪个来源能够解决不确定性。现有的主动式方法通常专注于用户澄清或环境验证中的一种,而没有明确地为每个决策确定合适的信息来源。我们将此问题形式化为在ACT(执行)、CLARIFY(澄清)和VERIFY(验证)之间进行不确定性路由,并提出了PROUR,一个主动式不确定性路由框架。PROUR将行动不确定性分解为:不同可能的用户目标解释之间的分歧(这标志着用户侧歧义),以及每个解释内部剩余的熵(这标志着缺失的世界侧证据)。为了从路由来源获取信息,我们训练了一个查询生成器,使用基于模式的条件信息增益奖励,在CLARIFY下针对用户目标识别,在VERIFY下针对下一步行动识别。在τ-bench上,PROUR在零售和航空领域实现了28.17%的平均成功率,比最强先前方法高出4.57%,同时使用的交互步骤减少了2.17步。学习到的策略进一步泛化到更强的任务智能体和τ³-bench的交易领域,无需重新训练,展示了来源对齐的不确定性解决对主动式智能体的益处。
英文摘要
Tool-using LLM agents must decide not only whether additional information is needed, but also which source can resolve the uncertainty. Existing proactive approaches often specialize in either user clarification or environment verification, without explicitly determining the appropriate information source for each decision. We formulate this problem as uncertainty routing among ACT, CLARIFY, and VERIFY, and propose PROUR, a proactive uncertainty routing framework. PROUR decomposes action uncertainty into disagreement across plausible user-goal interpretations, which signals user-side ambiguity, and the entropy remaining within each interpretation, which signals missing world-side evidence. To acquire information from the routed source, a query generator is trained with a mode-conditioned information-gain reward, targeting user-goal identification under CLARIFY and next-action identification under VERIFY. On $τ$-bench, PROUR achieves 28.17% average success rate across retail and airline, outperforming the strongest prior method by 4.57% while using 2.17 fewer interaction steps. The learned policy further generalizes to stronger task agents and transactional domains of $τ^3$-bench without retraining, demonstrating the benefit of source-aligned uncertainty resolution for proactive agents.