发表机构
East China Normal University; Shandong University(华东师范大学; 山东大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对检索增强大语言模型面临的证据选择问题,引入SelectBench基准和训练集,用DAPO对Qwen3.5 - 4B后训练,在测试集上提升了严格成功率,减少了禁止内容采用,但增益有限,还指出抗注入性和统计稳健性是后续挑战。
AI 中文摘要
检索增强的大语言模型常面临有用证据与误导性陈述或指令类内容交织的情况。直接拒绝会丢弃有效证据,不加批判地采用则会产生错误或不安全答案。因此,在现实检索场景中,有选择地采用相关信息同时拒绝欺骗性或有害内容的能力对可靠部署至关重要。我们引入了SelectBench,一个用于选择性证据采用的受控基准和训练集,并使用确定性规则奖励或冻结语义判断器通过DAPO直接对Qwen3.5 - 4B进行后训练。在修正后的325个示例的SelectBench - v2测试集上,原始检查点的严格成功率为22.46%,使用DAPO - Rule时升至25.54%,使用DAPO - DeepSeek时升至26.46%。两种训练策略都减少了对禁止内容的采用并产生了更简短、更有针对性的回答,但后续的提示注入并没有改善。配对增益不大且未通过霍尔姆校正,这表明可能需要更强的奖励塑造或更多的训练迭代来获得更稳健的增益。DAPO - DeepSeek在MMLU或干净的HotPotQA上没有明显退化,表明后训练过程保留了通用能力。这些结果表明在选择性证据使用方面有了方向性的改进,同时也指出了抗注入性和统计稳健性是未来工作中重要的剩余挑战。
英文摘要
Retrieval-augmented large language models frequently face contexts that interleave useful evidence with misleading statements or instruction-like content. Blanket refusal discards valid evidence, whereas uncritical adoption yields incorrect or unsafe answers. The ability to selectively adopt relevant information while rejecting deceptive or harmful content is therefore critical for reliable deployment in real-world retrieval settings. We introduce SelectBench, a controlled benchmark and training set for selective evidence adoption, and post-train Qwen3.5-4B directly with DAPO using either deterministic rule rewards or a frozen semantic judge. On the corrected 325-example SelectBench-v2 test set, strict success rises from 22.46% for the original checkpoint to 25.54% with DAPO-Rule and 26.46% with DAPO-DeepSeek. Both trained policies reduce forbidden-content adoption and produce shorter, more focused responses, yet prompt-injection following does not improve. The paired gains are modest and fail to survive Holm correction, suggesting that stronger reward shaping or additional training iterations may be needed for more robust gains. DAPO-DeepSeek exhibits no material degradation on MMLU or clean HotpotQA, indicating that the post-training procedure preserves general capabilities. These results demonstrate a directional improvement in selective evidence use, while identifying injection resistance and statistical robustness as important remaining challenges for future work.