arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

JKO-RAG:作为瓦瑟斯坦自由能梯度流的分布检索

JKO-RAG: Distributional Retrieval as Wasserstein Free-Energy Gradient Flow

Levi Segal, Murari Ambati

arXiv 2607.24776首次发表:更新:

AI 中文总结

研究针对RAG管道排名列表与下游语言模型集合条件设定的不匹配问题,提出JKO方法,将重新排序视为在瓦瑟斯坦-2梯度流下最小化自由能泛函,通过线性响应理论解释优势,经实验验证,扩展方法并在BEIR基准测试中表现出色。

AI 中文摘要

RAG管道返回一个段落的排名列表。我们认为这存在不匹配:下游语言模型基于集合进行条件设定,而选择问题本质上是几何问题。我们提出了JKO,它将重新排序框架化为在瓦瑟斯坦-2梯度流下通过约旦 - 金德勒勒 - 奥托近端方案最小化自由能泛函\(F(p)=\text{相关性}+\text{熵}+\text{冗余}\)。地面度量\(C_{ij}=(1 - \cos\langle z_i,z_j\rangle)^2\)编码了嵌入流形的语义几何。我们的核心贡献是一种线性响应理论,解释了为什么瓦瑟斯坦几何有帮助:瓦瑟斯坦和KL检索映射仅在其近端海森矩阵上不同——对于\(W^2\)是密集且几何感知的,对于KL是对角且几何盲的——这种差异抑制了查询释义引起的质量传输。该理论产生了一个可证伪的预测:稳定性优势在步长\(h\)中单调递减。我们通过自由能下降、频率分辨扰动响应、预测的\(h\)依赖性和认证半径分析进行了实证验证。引入了四个扩展:\(\textbf{\nmjko}\)(学习的地面度量)、\(\textbf{\bwjko}\)(\(W^2\) - KL插值)、\(\textbf{\samjko}\)(加速两倍)和\(\textbf{\dualrank}\)(OT对偶势作为置信信号)。在五个BEIR基准测试中,JKO在所有五个测试中都优于交叉编码器;决定性优势是鲁棒性——在释义下稳定22 - 38%,泄漏的干扰物减少两倍。

英文摘要

RAG pipelines return a \emph{ranked list} of passages. We argue this is a mismatch: the downstream language model conditions on a \emph{set}, and the selection problem is fundamentally geometric. We propose \jko, which frames reranking as minimising a free-energy functional $F(p)=\text{relevance}+\text{entropy}+\text{redundancy}$ under Wasserstein-2 gradient flow via the Jordan--Kinderlehrer--Otto proximal scheme. The ground metric $C_{ij}=(1-\cos\langle z_i,z_j\rangle)^2$ encodes the semantic geometry of the embedding manifold. Our central contribution is a \emph{linear-response theory} explaining \emph{why} the Wasserstein geometry helps: the Wasserstein and KL retrieval maps differ only in their proximal Hessian -- dense and geometry-aware for $W^2$, diagonal and geometry-blind for KL -- and this difference damps the mass transport that query paraphrase induces. The theory yields a falsifiable prediction: the stability advantage is monotonically decreasing in step size $h$. We verify this empirically via free-energy descent, frequency-resolved perturbation response, the predicted $h$-dependence, and a certified-radius analysis. Four extensions are introduced: \textbf{\nmjko} (learned ground metric), \textbf{\bwjko} ($W^2$--KL interpolation), \textbf{\samjko} ($2\times$ speedup), and \textbf{\dualrank} (OT dual potentials as confidence signals). Across five BEIR benchmarks, \jko\ outperforms the cross-encoder on all five; the decisive advantage is robustness -- 22--38\% more stable under paraphrase, $2\times$ fewer leaked distractors.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑