你的检索器已经知道:用于RAG检索充分性的分布形状QPP
Your Retriever Already Knows: Distribution-Shape QPP for RAG Retrieval Sufficiency
浏览论文内容
中文总结 AI 辅助
本文提出分布形状QPP特征用于RAG检索充分性预测,在ViDoRe和SÚJB上达到高AUROC,速度快于LLM评判器,且混合方法在对抗检测上更优。
中文摘要 AI 辅助
标准的检索增强生成(RAG)流程通常无法在推理时提供可靠的信号来判断检索是否成功;对于模糊或超出范围的问题,生成过程可能会产生幻觉。受捷克核监管机构部署场景的启发,该场景中数据敏感性排除了第三方LLM API的使用,我们比较了三种用于RAG检索充分性的查询性能预测(QPP)范式:基于分数的特征、基于内容的LLM评判器以及混合方法。在八个ViDoRe视觉领域(14,514个查询)上,我们的24个非词汇特征(GeneralQPP;15个分布形状、5个查询表面、4个全局)在每查询2毫秒内达到了0.856的加权平均AUROC,优于经典QPP文献池(Classic Full,0.835),并远高于本地Qwen3.5 LLM评判器(0.649,差距+0.207;每查询快约3000倍且更便宜)。将LLM评判作为特征之一(混合方法)在ViDoRe上匹配S1(0.863),但在SÚJB上获得了统计显著的提升(AUROC 0.911,Hit@5,对抗检测0.954;1,510个查询,500个合成对抗样本),但需承担LLM延迟。排名在不同数据集间一致(Spearman ρ = 0.90)。在留一域外验证下,S1降至0.706;一个13特征的LODO逐步子集(S1-Lean)恢复至0.719(比文献池高+0.032)。
英文摘要
Standard Retrieval-Augmented Generation (RAG) pipelines often provide no reliable inference-time signal of whether retrieval succeeded; on ambiguous or out-of-scope queries, generation may then hallucinate. Motivated by a Czech nuclear-regulator deployment where data sensitivity precludes third-party LLM APIs, we compare three Query Performance Prediction (QPP) paradigms for retrieval sufficiency in RAG: score-based features, a content-based LLM judge, and a hybrid. On the eight ViDoRe vision domains (14,514 queries), our 24 non-lexical features (GeneralQPP; 15 distribution-shape, 5 query-surface, 4 global) reach a weighted-average AUROC of 0.856 at 2 ms per query, ahead of a classic-QPP literature pool (Classic Full, 0.835) and well above a local Qwen3.5 LLM judge (0.649, +0.207 gap; $\sim$3000$\times$ faster and cheaper per query). Adding the LLM judgment as one feature (hybrid) matches S1 on ViDoRe (0.863) but gains a statistically significant edge on SÚJB (AUROC 0.911 at Hit@5, adversarial-detection 0.954; 1,510 queries, 500 synthetic adversarial), at LLM latency. Rankings agree across datasets (Spearman $ρ= 0.90$). Under Leave-One-Domain-Out, S1 drops to 0.706; a 13-feature LODO-stepwise subset (S1-Lean) recovers to 0.719 (+0.032 over the literature pool).
发表机构
- Department of Mathematics, FNSPE CTU in Prague(查尔斯理工大学应用科学与工程学院数学系)
- Institute of Information Theory and Automation, Czech Academy of Sciences(捷克科学院信息与自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。