回答还是弃权:通过弃权感知强化学习减轻搜索代理幻觉
To Answer or to Abstain: Mitigating Search-Agent Hallucinations via Abstention-Aware Reinforcement Learning
AI总结:
研究针对大语言模型训练范式易加剧搜索代理幻觉的问题,提出弃权感知强化学习(AWA-RL),利用模型能力动态塑造弃权奖励,并引入RA-F1指标,实验表明该方法有效提升了搜索代理的性能和可靠性。
AI中文摘要:
近期在为大语言模型配备搜索工具和结果奖励强化学习方面的进展,在开放域问答任务上取得了新的最优结果。然而,当前训练范式存在关键漏洞:主要奖励正确答案,检索失败时却不惩罚编造答案,从而加剧幻觉。为此提出弃权感知强化学习(AWA-RL),利用模型特定查询先验能力和持续在线策略训练观测动态塑造弃权奖励。还引入新指标RA-F1衡量能力-可靠性权衡。与非弃权基线相比,AWA-RL绝对精度提高多达10.3%,整体RA-F1提高2.9%,原始准确率仅有轻微牺牲。结果证实AWA-RL成功产生了高性能且可靠的搜索代理。
英文摘要:
Recent advances in equipping Large Language Models (LLMs) with search tools and outcome-reward reinforcement learning (RL) have achieved new state-of-the-art results on open-domain QA tasks. However, we argue that current training paradigms harbor a critical vulnerability: they predominantly reward correct answers but fail to penalize fabricated ones when retrieval fails, thereby implicitly exacerbating hallucinations. To address this, we propose Abstention-Aware Reinforcement Learning (AWA-RL), which dynamically shapes the abstention reward utilizing the model's query-specific prior capabilities and continuous on-policy training observations. We also introduce a novel metric, RA-F1, to measure the capability-reliability trade-off. Compared to non-abstaining baselines, AWA-RL boosts absolute precision by up to 10.3% and overall RA-F1 by 2.9%, with only marginal sacrifice in raw accuracy. These results confirm that AWA-RL successfully yields highly capable and reliable search agents. The code, data, and model weights are publicly available at https://github.com/zfj1998/AWA-RL.