arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2606.30518cs.CL

基于体制感知的同行专业化方法,用于在异构知识冲突下实现稳健的检索增强生成

Regime-Aware Peer Specialization for Robust RAG under Heterogeneous Knowledge Conflicts

  • School of Computer Science, Beijing Institute of Technology(北京理工大学计算机科学学院)

机构由 AI 辅助整理,请以论文原文为准。

Bo Wang, Heyan Huang, Yaolin Li, Yanghao Zhou, Jiahao Teng, Ziyi Yang, Ge Shi, Chong Feng

更新

AI总结:

提出RAPS-DA框架,通过样本级三体制划分(接地、仲裁、抵抗)和token级双选择器,解决RAG中异构知识冲突问题,在固定模型规模下超越多种基线。

AI中文摘要:

检索增强生成(RAG)通过将生成过程锚定在外部上下文中来改进语言模型。然而,当检索到的上下文与模型的参数化知识冲突时,它可能变得脆弱。这种冲突涵盖了一个可靠性谱系,从可靠和部分可靠的证据到对抗性上下文。现有的补救措施通常采用体制无关的监督来处理这种异构冲突,这可能会混淆不同可靠性体制下的不兼容学习信号。为了解耦这些信号,我们提出了RAPS-DA,一个体制感知的同行专业化框架,在两个互补的粒度上解决冲突。在样本级别,冲突被分为三个体制,包括接地、仲裁和抵抗,每个体制从共享的基础模型训练一个同尺度的同行专家。然后,每个样本被硬路由到其体制匹配的同行,进行策略上的反向KL监督。在token级别,一个双层选择器使用教师间分歧、学生-教师差异和学生熵来过滤无信息或不稳定的token,对自信不匹配的token进行上加权,并随着学生的成熟逐渐将监督集中在高冲突token上。收益来自于固定模型尺度上的专业化,而不是更强的教师,并且同行专家仅在训练期间存在,因此部署的学生不需要体制标签或同行访问。在五个冲突场景和两个分布外基准上的实验表明,RAPS-DA超越了所有提示、解码、微调、强化学习和单教师基线。

英文摘要:

Retrieval-augmented generation (RAG) improves language models by grounding generation in external context. However, it can be fragile when the retrieved context conflicts with the model's parametric knowledge. Such conflicts span a reliability spectrum, ranging from reliable and partially reliable evidence to adversarial context. Existing remedies often handle such heterogeneous conflicts with regime-agnostic supervision, which can conflate incompatible learning signals across reliability regimes. To disentangle these signals, we propose RAPS-DA, a regime-aware peer specialization framework that addresses conflict at two complementary granularities. At the sample level, conflicts are divided into three regimes, including Grounding, Arbitration, and Resistance, with one same-scale peer specialist trained per regime from a shared base model. Each sample is then hard-routed to its regime-matched peer for on-policy reverse-KL supervision. At the token level, a dual-layer selector uses inter-teacher disagreement, student-teacher divergence, and student entropy to filter uninformative or unstable tokens, upweight confidently misaligned ones, and gradually focus supervision on high-conflict tokens as the student matures. Gains stem from specialization at a fixed model scale, not from a stronger teacher, and the peer specialists exist only during training, so the deployed student requires no regime labels or peer access. Experiments on five conflict scenarios and two out-of-distribution benchmarks show RAPS-DA surpasses all prompting, decoding, fine-tuning, RL, and single-teacher baselines.

补充信息

↑