arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HERALD:检索证明奖励的反事实审计与最小修复

HERALD: Counterfactual Audits and Minimal Repairs for Proof-of-Retrieval Rewards

Zhuowen Liu, Bohan Cui, YinShang Guo, Yuting Wang, Hao Li

arXiv 2608.06012首次发表:更新:

AI 中文总结

HERALD是一种离线审计方法,可对搜索智能体的检索证明奖励进行反事实审计与最小修复,能有效抵御引用清洗攻击,提升引用精度等指标并分离稳健评分与策略迁移。

AI 中文摘要

搜索智能体的奖励包含答案质量、引用依据、工具成本和反黑客条款;因此高分不一定意味着检索到了引用证据,且额外惩罚可抵消高分。本文提出HERALD,一种离线审计方法,应用完全相同问题的干预措施,区分候选可见信息与预言机信息,并在策略优化前枚举检测器契约。在来自HotpotQA、2WikiMultiHopQA和MuSiQue的四个Qwen3-8B池上,R₀拒绝搜索删除和虚假ID,但无标签的引用清洗攻击仍能成功。完整的2³消融实验确定,针对性增强L(即引用检索证据中不存在的语篇段落)是观察到的包含最小修复:R[L]的经验ASR为零,单侧集群上界为0.50%。该差距在池规则、可见BM25攻击者和四个模型中持续存在;当攻击移除预言机支持ID惩罚时,更广泛的加固仍易受攻击。在严格的5M token匹配训练(每个基准评估256对问题)下,R[L]在HotpotQA和2Wiki上达到EM非劣性阈值,但在MuSiQue上未达到。同组的引用精度和支持召回率分别提升2.02和1.46个百分点,无支持引用下降1.69个百分点,2Wiki和MuSiQue上的清洗攻击可降低。自然L未减少,且检测器仅出现在58368条训练轨迹中的18条。因此HERALD实现了稳健评分、稀疏学习信号和策略迁移的分离。

英文摘要

Search-agent rewards mix answer quality, citation grounding, tool cost, and anti-hacking terms; a high score therefore need not imply that cited evidence was retrieved, and added penalties can cancel. We introduce HERALD, an offline audit that applies exact same-question interventions, separates candidate-visible from oracle information, and enumerates detector contracts before policy optimization. On four Qwen3-8B pools from HotpotQA, 2WikiMultiHopQA, and MuSiQue, $R_0$ rejects search deletion and fake IDs, but a label-free citation-laundering attack succeeds. A complete $2^3$ ablation identifies targeted strengthening of $L$---citing a corpus passage absent from the retrieved evidence---as the observed inclusion-minimal repair: $R[L]$ has zero empirical ASR with a 0.50% one-sided cluster upper bound. The gap persists across pool rules, a visible BM25 attacker, and four models; broader hardening remains vulnerable when the attack removes an oracle support-ID penalty. Under strict 5M-token matched training evaluated on 256 paired questions per benchmark, $R[L]$ meets the EM non-inferiority gate on HotpotQA and 2Wiki but not MuSiQue. Equal-suite citation precision and support recall improve by 2.02 and 1.46 points, unsupported citations fall by 1.69, and laundering attackability falls on 2Wiki and MuSiQue. Natural $L$ is not reduced, and the detector appears in only 18 of 58,368 training trajectories. HERALD thus separates robust scoring, sparse learning signal, and policy transfer.

Comments9 pages, 3 figures, and 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑