用于检索增强生成的源感知重排:一种可靠性先验方法
Source-Aware Reranking for Retrieval-Augmented Generation: A Reliability Prior Approach
浏览论文内容
中文总结 AI 辅助
研究针对检索增强生成中仅靠语义相似性排序的问题,提出纳入源可靠性先验的重排方法,通过为文档分配先验并重新加权检索分数,在健康领域语料库实验中提升了Precision@5并减少对抗性文档检索。
中文摘要 AI 辅助
标准的检索增强生成管道仅通过语义相似性对检索到的文档进行排序,而不考虑源出处或可信度。这项工作评估了对RAG检索排名的一种简单且可解释的修改,该修改纳入了领域相关的源可靠性先验。根据文档的源类型为每个文档分配一个先验lambda(s),并使用score(q, d) = sim(q, d) * lambda(s)对检索分数进行重新加权。在一个120篇文档的健康领域语料库上,将该框架与仅基于相似性的基线进行评估。在这种受控设置下,源感知重排将Precision@5从0.48提高到0.72,并在评估的威胁模型下减少了平均对抗性文档检索,其中低可信度源可通过元数据识别。所有实验都在密尔沃基工程学院的高性能计算集群Rosie上执行,该集群提供了可靠且可重复运行完整实验管道所需的GPU加速基础设施。这些结果表明,在所述实验设置的范围内,RAG管道中源质量下降的潜在缓解策略。
英文摘要
Standard Retrieval-Augmented Generation pipelines rank retrieved documents by semantic similarity alone, without accounting for source provenance or credibility. This work evaluates a simple and interpretable modification to RAG retrieval ranking that incorporates domain-informed source reliability priors. Each document is assigned a prior lambda(s) based on its source type, and retrieval scores are reweighted using score(q, d) = sim(q, d) * lambda(s). The framework is evaluated against a similarity-only baseline on a 120-document health-domain corpus. In this controlled setting, source-aware reranking improves Precision@5 from 0.48 to 0.72 and reduces average adversarial document retrieval under the evaluated threat model, where low-credibility sources are identifiable via metadata. All experiments were executed on Rosie, the high-performance computing cluster at the Milwaukee School of Engineering, which provided the GPU-accelerated infrastructure necessary to run the full experimental pipeline reliably and reproducibly. These results suggest a potential mitigation strategy for source quality degradation in RAG pipelines, within the limits of the experimental setup described.
发表机构
- Milwaukee School of Engineering(密尔沃基工程学院)
机构由 AI 辅助整理,请以论文原文为准。