arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Socrates-RAG:针对协同证据投毒的基于前提的主动检索

Socrates-RAG: Premise-Directed Inquiry against Coordinated Evidence Poisoning

Renyu Zhao, Xinyuan Zou, Lanbin Liu

arXiv 2609.35773首次发表:更新:

发表机构

Tencent SSV; School of Future Cities, University of Science and Technology Beijing(腾讯SSV; 北京科技大学未来城市学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Socrates-RAG提出基于前提的主动检索策略,通过选择未解决前提并优化查询,在开放语料库问答中对抗证据投毒,将安全准确率从79.2%提升至93.8%。

AI 中文摘要

检索增强生成(RAG)的防御机制通常决定如何过滤或聚合一个固定的检索集合。然而,在开放语料库问答中,决定性的证据可能不在初始上下文中,但可以通过检索获得,这使得下一次查询成为可靠性问题的一部分。我们引入了Socrates-RAG,一种基于前提的主动检索策略,该策略表示相互竞争的答案,选择一个其解决能够区分这些答案的未解决前提,并在回答或弃权(不执行)之前利用新获取的证据来优化后续查询。我们形式化了由此产生的有限预算证据状态,并给出了相对于重复或主题查询策略的条件性救援保证。我们将Socrates-RAG与一个匹配的对照组进行评估,在该对照组中,相同的骨干模型生成普通的面向相关性的搜索查询;两种策略共享初始证据、确定性检索器、两次查询/前三项预算、答案提示和无标签证据链发布规则。在不相交的48个世界反事实评估中,基于前提的查询将安全准确率从79.2%提升至93.8%,其中8胜1负39平(双侧精确p=.0391)。决定性证据召回率以相同幅度提高,而不安全答案从1个降至0个。两种策略都解决了所有24个单跳案例;收益集中在两跳案例中,其中Socrates-RAG将新解决的前提替换到其第二次查询中。这项受控研究隔离了基于前提获取的特定益处,而不声称在开放网络上具有普遍鲁棒性。

英文摘要

Retrieval-augmented generation (RAG) defenses typically decide how to filter or aggregate a fixed retrieved set. In open-corpus question answering, however, decisive evidence may be absent from the initial context but retrievable, making the next query part of the reliability problem. We introduce Socrates-RAG, a premise-directed active retrieval policy that represents competing answers, selects an unresolved premise whose resolution would discriminate them, and uses newly acquired evidence to refine a subsequent query before answering or abstaining. We formalize the resulting finite-budget evidence state and give a conditional rescue guarantee relative to repeated or topical-query policies. We evaluate Socrates-RAG against a matched control in which the same backbone generates ordinary relevance-oriented search queries; both policies share the initial evidence, deterministic retriever, two-query/top-three budget, answer prompt, and label-free evidence-chain release rule. On a disjoint 48-world counterfactual evaluation, premise-directed inquiry raises safe accuracy from 79.2% to 93.8%, with 8 wins, 1 loss, and 39 ties (two-sided exact p=.0391). Decisive-evidence recall improves by the same margin, while unsafe answers fall from one to zero. Both policies solve all 24 one-hop cases; the gain is concentrated in two-hop cases, where Socrates-RAG substitutes a newly resolved premise into its second query. This controlled study isolates a specific benefit of premise-directed acquisition without claiming general robustness on the open Web.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑