arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02009cs.AI

HALT:面向检索增强搜索智能体的感知验证式停止策略

HALT: Verification-Aware Stopping for Retrieval-Augmented Search Agents

  • KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

Daeyoung Roh, Donghee Han

AI总结:

该研究针对检索增强搜索智能体的冗余检索问题,提出轻量级感知验证停止策略HALT,在三个多跳QA基准上减少冗余搜索且保持精确匹配,无需修改宿主智能体。

AI中文摘要:

检索增强搜索智能体通过反复发出搜索查询并积累证据来回答多跳问题,这会产生停止问题:在必要证据出现后,进一步检索往往会增加成本、延迟以及分散注意力的上下文,而非有用信息。我们将停止问题建模为证据覆盖问题,而非生成器置信度问题,并引入HALT,这是一种轻量级的感知验证策略,可保持搜索智能体不变。给定预期的跳数主张,HALT仅在累积证据支持每个所需主张时停止。在三个多跳QA基准测试中,HALT减少了冗余搜索,同时在很大程度上保持了精确匹配。我们区分了可部署设置(其中跳数主张由问题生成)和使用黄金支持事实注释的诊断上限:生成的主张带来较小但仍保持精确匹配的节省,而黄金主张则显示了当跳数目标清晰时可获得的更大节省。基线比较和消融实验表明,这种行为由主张-证据对齐驱动,而非通用充分性、固定停止位置或词汇重叠。开放语料库试点进一步表明,当无法可靠验证覆盖范围时,HALT会弃权(不执行)。总体而言,证据覆盖为改进检索增强智能体提供了实用的运行时控制信号,无需重新训练或修改宿主智能体。

英文摘要:

Retrieval-augmented search agents answer multi-hop questions by repeatedly issuing search queries and accumulating evidence. This creates a stopping problem: after the necessary evidence has appeared, further retrieval often adds cost, latency, and distracting context rather than useful information. We frame stopping as evidence coverage rather than generator confidence, and introduce HALT, a lightweight verification-aware policy that leaves the search agent unchanged. Given expected hop claims, HALT stops only when cumulative evidence supports each required claim. Across three multi-hop QA benchmarks, HALT reduces redundant search while largely preserving exact match. We separate a deployable setting, where hop claims are generated from the question, from a diagnostic upper bound that uses gold supporting-fact annotations: generated claims give smaller but still exact-match-preserving savings, while gold claims show the larger savings available when hop targets are clean. Baseline comparisons and ablations show that this behavior is driven by claim-evidence alignment rather than generic sufficiency, fixed stop positions, or lexical overlap. Open-corpus pilots further suggest that HALT abstains when coverage cannot be reliably verified. Overall, evidence coverage provides a practical runtime control signal for improving retrieval-augmented agents without retraining or modifying the host agent.

补充信息

↑