arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RIMS:通过平滑多对聚合进行偏好优化以实现小规模语言模型检索增强生成

RIMS: Preference Optimization via Smoothed Multi-pair Aggregation for Small-Scale LLM Retrieval-Augmented Generation

Pei Tian, Zihan Dong, Tianci Liu, Linjun Zhang, Haoyu Wang

arXiv 2607.16431首次发表:更新:

发表机构

Columbia University; Rutgers University; Purdue University; SUNY Albany(哥伦比亚大学; 罗格斯大学; 普渡大学; 纽约州立大学奥尔巴尼分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对小规模语言模型检索增强生成中对噪声证据敏感的问题,提出RIMS框架,通过合成偏好数据、软聚合机制及偏好优化,实现更好性能,在多基准测试中优于现有方法

AI 中文摘要

小规模语言模型(SLMs)在资源受限环境下对检索增强生成(RAG)很有吸引力,但容量有限使其对噪声或虚假检索证据高度敏感。现有基于偏好的方法存在信号丢弃或数据利用率低的问题。我们提出RIMS,一个三阶段偏好优化框架,包括利用目标SLM本身通过拒绝采样生成合成思维链偏好数据、用平滑算子替代硬选择的可微软聚合机制、对多种对齐算法应用平滑目标进行偏好优化。理论表明平滑近似有可控误差界且软聚合梯度对齐更紧密。实验显示该方法在多个SLM骨干上优于现有基线。

英文摘要

Small-scale language models (SLMs) are attractive for retrieval-augmented generation (RAG) in resource-constrained settings, but their limited capacity makes them highly sensitive to noisy or spurious retrieved evidence. Existing preference-based methods such as RoseRAG select only the hardest single preference pair via hard argmin/argmax, discarding the remaining signal; others treat multiple pairs as independent binary comparisons, resulting in low data utilization. We propose RIMS, a three-stage preference optimization framework comprising (1) synthetic chain-of-thought preference data generation via rejection sampling using the target SLM itself without relying on proprietary models, (2) a differentiable soft aggregation mechanism that replaces hard selection with a smooth operator, preserving gradient signal from all preference pairs while retaining the discriminative structure of margin-aware selection, and (3) preference optimization with the smoothed objective applied to multiple alignment algorithms. We theoretically show that the smoothed approximation admits a controllable error bound and that smooth aggregation yields provably tighter gradient alignment to the oracle objective than hard selection. Experiments on four multi-hop question answering benchmarks show that our approach outperforms state-of-the-art baselines across multiple SLM backbones, achieving consistent gains in Exact Match and F1 under noisy retrieval conditions. Our implementation is available at https://github.com/tptrix29/RIMS.

Journal refCOLM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑