发表机构
UCAS; Renmin University of China(中国科学院大学; 中国人民大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究深度搜索中委托智能,将其分解为搜索决策和信息合成与验证维度。通过开发可控合成管道构建DelegSearchBench及评估协议,实验证明仅用最终答案准确性不能充分表征深度搜索能力。
AI 中文摘要
深度搜索正成为现代智能体系统的核心能力,但通常仅基于端到端答案准确性进行评估。这种耦合评估范式将检索质量、长上下文理解、证据验证和工具使用决策纠缠在一起,难以确定模型是否真正知道何时以及如何委托信息搜索。为此,我们将这种元能力形式化为深度搜索中的委托智能,并将其分解为互补维度——搜索决策(识别信息不足并决定是否、何时以及如何搜索)和信息合成与验证(从多个来源聚合证据、判断源可靠性并在有噪声、潜在对抗条件下合成信息)。为实现解缠且可重复的测量,我们开发了基于文档逆向工程的可控合成管道,构建了DelegSearchBench及解缠评估协议。通过实验表明,仅靠最终答案准确性无法充分表征深度搜索能力。
英文摘要
Deep search is becoming a core capability of modern agent systems, yet it is typically evaluated solely based on end-to-end answer accuracy. This coupled evaluation paradigm entangles retrieval quality, long-context comprehension, evidence verification, and tool-use decisions, making it difficult to determine whether a model truly knows when and how to delegate information seeking to search. To this end: (1) We formalize this meta-capability as Delegation Intelligence in deep search and decompose it into complementary dimensions-Search Decision-Making (recognizing information insufficiency and deciding whether, when, and how to search) and Information Synthesis and Verification (aggregating evidence from multiple sources, judging source reliability, and synthesizing information under noisy, potentially adversarial conditions). (2) To enable disentangled and reproducible measurement, we develop a controllable synthesis pipeline built on document-grounded reverse engineering. This yields a general recipe for constructing controlled deep-search evaluations rather than a single fixed dataset. (3) As a concrete instantiation, we construct DelegSearchBench, together with a disentangled evaluation protocol that isolates each capability dimension by varying document composition and tool access. (4) Across representative models, we demonstrate that deep-search competence cannot be adequately characterized by final-answer accuracy alone...
CommentsWork in Progress