发表机构
Ant Group; NLPR, MAIS, CASIA(蚂蚁集团; 中国科学院自动化研究所模式识别国家重点实验室、多模态人工智能系统实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对多跳搜索增强智能体的弃权问题,构建首个可控基准HopRefusalBench,经实验发现前沿模型弃权能力存在显著不足,为相关可靠性改进提供基础。
AI 中文摘要
搜索增强型大语言模型智能体日益具备解决知识密集型任务的能力,但当多跳问题根本无法回答时,其行为模式仍未得到充分理解。现有的弃权基准大多仅能暴露单跳查询层面的缺陷,无法揭示经有效中间推理与检索后才会出现的失败情况。我们推出HopRefusalBench,这是首个针对多跳搜索内弃权的可控基准,包含889个基于KILT关联实体路径构建的无法回答问题。它涵盖三类无法回答的原因(答案未知、前提错误、上下文未明确),以及根、中间、终端三种拓扑结构,可分别观测前提验证、中间桥梁验证和终端停止的情况。我们还提出了最终结果分类体系,涵盖目标感知弃权、伪弃权、幻觉补全、搜索预算耗尽,以及用于触发后延续和token浪费的源感知轨迹指标。在10个采用搜索增强模式的前沿专有与开源模型中,最优模型的目标感知正确停止率(TCHR)仅为42.9%。根和中间类问题始终比终端类问题更难,所有模型在错误前提问题上的TCHR最高,在未明确上下文问题上的TCHR最低。但按类别汇总后,每个模型84.7%至98.4%的明确弃权类响应都能识别正确理由,主要瓶颈在于确定合适的非答案;失败轨迹则会分化为幻觉或搜索预算耗尽。这些结果表明,多跳搜索中的弃权是一个重要的评估问题,并为诊断和改进搜索增强智能体的可靠性提供了基础。
英文摘要
Search-augmented large language model agents are increasingly capable of solving knowledge-intensive tasks, but their behavior when a multi-hop question is fundamentally unanswerable remains poorly understood. Existing abstention benchmarks largely expose defects at the surface of single-hop queries and therefore cannot reveal failures that emerge only after valid intermediate reasoning and retrieval. We introduce HopRefusalBench, the first controlled benchmark of refusal within multi-hop search, comprising 889 unanswerable questions constructed from KILT-grounded entity paths. It crosses three causes of unanswerability (answer unknown, false premise, and underspecified context) with root, middle, and terminal topologies, making premise verification, intermediate-bridge validation, and terminal stopping separately observable. We further propose a final-outcome taxonomy spanning target-aware refusal, pseudo-refusal, hallucinated completion, and search-budget exhaustion, together with source-aware trajectory metrics for post-trigger continuation and token waste. Across ten frontier proprietary and open-weight models in search-augmented mode, the best model achieves a target-aware correct halting rate (TCHR) of only 42.9%. Root and middle items are consistently harder than terminal items, and all models attain their highest TCHR on false premises and their lowest on underspecified questions. Yet when pooled across categories, 84.7--98.4% of each model's explicit refusal-like responses identify the correct rationale, localizing the main bottleneck to committing to an appropriate non-answer; failed trajectories instead diverge into hallucination or search-budget exhaustion. These results establish refusal in multi-hop search as a consequential evaluation problem and provide a foundation for diagnosing and improving the reliability of search-augmented agents.
Comments20 pages