arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

询问未被请求的内容:智能体中的水平与垂直主动性

Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents

Ido Levy, Asaf Yehudai, Segev Shlomov, Asaf Adi, Leshem Choshen

arXiv 2609.37236首次发表:更新:

发表机构

IBM; Weizmann Institute of Science(IBM; 魏茨曼科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究智能体主动性的内容维度,提出水平与垂直主动性,并设计Q&D方法训练提问者,在无需奖励模型下提升多跳问答中的信息获取,优于15倍大模型。

AI 中文摘要

使用工具智能体通常响应于用户的明确请求,但完成任务可能需要用户从未请求的信息。关于主动智能体的研究主要关注智能体是否以及何时应自主行动,而非应追求什么信息。我们研究了主动性的一个不同维度:其内容。水平主动性追求当前上下文已识别的未陈述信息,而垂直主动性追求仅由早期证据揭示的需求。从基准自身的分解中恢复的需求图记录了哪些需求依赖于哪些需求,因此两种形式以及智能体是否在正确时间停止,都可以从转录中评分,无需模型评判器。为了学习这种行为,我们提出了Q&D(提问者和起草者),它训练一个提问者偏好于其续篇检索更多所需证据的问题,无需奖励模型或评判器。在三个多跳问答基准的保留分割上,在相同的检索开销下,经过训练的提问者相对于相同模型的提示版本提高了两种形式的主动性,并在三个基准中的两个上优于相同角色中提示的15倍大模型,且在控制问题数量和长度后增益持续存在。无需进一步训练,我们将提问者置于一个与模拟客户互动的交互式客户服务智能体中,在那里它以更少的问题完成更多任务,在零售领域,它以更少的客户后续轮次优于15倍大模型。这些结果表明,主动性不仅取决于智能体是否在未被要求时行动,还取决于它选择追求什么以及何时停止。

英文摘要

An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on its own, not what information it should pursue. We study a distinct axis of proactivity: its content. Horizontal proactivity pursues unstated information that the current context already identifies, and vertical proactivity pursues needs that only earlier evidence reveals. A need graph, recovered from a benchmark's own decomposition, records which needs depend on which, so both forms, and whether the agent stops at the right time, can be scored from a transcript without a model judge. To learn this behavior, we propose Q&D (questioner and drafter), which trains a questioner to prefer the question whose continuation retrieves more of the required evidence, with no reward model or judge. On held-out splits of three multi-hop question-answering benchmarks, at equal retrieval spend, the trained questioner improves both forms of proactivity over the same model, prompted, and outperforms a prompted model $15\times$ larger in the same role on two of the three, and the gain persists after controlling for question volume and length. Without further training, we place the questioner in an interactive customer-service agent with a simulated customer, where it completes more tasks while asking fewer questions, and in retail it outperforms the $15\times$ larger model with fewer follow-up turns from the customer. These results show that proactivity depends not only on whether an agent acts without being asked, but also on what it chooses to pursue and when it stops.

Comments48 pages. Project page: https://dolev31.github.io/ProactiveInquirer/ Code: https://github.com/dolev31/ProactiveInquirer Model: https://huggingface.co/dolev31/ProactiveInquirer-Qwen3-8B

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑