从惯性到客观性:通过噪声隔离改进深度研究智能体
From Inertia to Objectivity: Improving Deep Research Agents with Noise Isolation
- School of Software Technology, Zhejiang University(浙江大学软件学院)
- Zhejiang University(浙江大学)
- Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对深度研究智能体存在的惯性偏差问题,提出NIS-Agent方法,在网页分类和最终答案验证环节应用上下文隔离,降低33%的token成本,训练的8B模型性能可与GPT-4o相当。
AI中文摘要:
由大语言模型(LLM)驱动的网页搜索智能体展现出强大潜力,但深度研究任务暴露出一种反复出现的失效模式:一旦智能体生成了查询、计划或中间结论,后续判断该相同行动的结果时就会变得不那么客观。我们将此现象称为惯性偏差(inertia bias)。为了对其进行量化,我们引入了IBIS基准,该基准在控制搜索观测结果的同时,改变模型是否在评估自身先前行动的结果。我们发现,当模型“拥有”前一个搜索步骤时,其性能会显著下降,表明自主生成的行动历史会系统性地扭曲后续判断。我们进一步表明,这种偏差会传播为两种系统级退化形式:工作者层面的搜索噪声和管理者层面的上下文噪声。为解决该问题,我们提出了NIS-Agent,该方法在最易受惯性偏差影响的两个决策点(网页分类和最终答案验证)应用上下文隔离。在GAIA、WebWalkerQA、BrowseComp和BrowseComp-zh数据集上,NIS-Agent实现了具有竞争力的性能,同时与基线相比降低了33%的token成本。我们进一步训练了一个8B模型,使其内在地更能抵抗惯性偏差;在相同的NIS-Agent框架下,该模型在深度研究基准上的平均性能可与GPT-4o相当。
英文摘要:
Web search agents powered by Large Language Models (LLMs) show strong promise, but deep research tasks expose a recurring failure mode: once an agent has produced a query, plan, or intermediate conclusion, it becomes less objective when later judging the consequences of that same action. We term this phenomenon inertia bias. To make it measurable, we introduce the IBIS benchmark, which controls the search observations while varying whether the model is evaluating the outcome of its own prior action. We find that models are substantially worse when they "own" the preceding search step, showing that self-authored action history can systematically distort subsequent judgment. We further show that this bias propagates into two forms of system-level degradation: search noise at the worker level and contextual noise at the manager level. To address this problem, we propose NIS-Agent, which applies context isolation at the two decision points most vulnerable to inertia bias: webpage triage and final-answer validation. Across GAIA, WebWalkerQA, BrowseComp, and BrowseComp-zh, NIS-Agent achieves competitive performance while reducing token cost by 33% compared to our baseline. We further train an 8B model to be intrinsically more resistant to inertia bias; under the same NIS-Agent framework, it attains average performance comparable to GPT-4o on deep research benchmarks. Our code is publicly available at https://github.com/PangSMPang/NIS-Agent.