FARCA:面向具备事实监督的强化学习的事实对齐可靠性感知信用分配
FARCA: Fact-Aligned Reliability-Aware Credit Assignment for Reinforcement Learning with Factual Supervision
- School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院)
- School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
FARCA是一种面向具备事实监督的强化学习的策略优化框架,通过对齐事实验证与策略更新的粒度、引入反事实证据归因计算可靠性权重,提升模型事实性并保留通用推理能力。
AI中文摘要:
为降低通过可验证奖励的强化学习训练的大型语言模型中由结果驱动奖励导致的幻觉风险,现有缓解方法引入了过程级事实监督。然而,由于事实信号的粗粒度聚合以及对这些信号的可靠性评估缺失,它们造成了事实验证与策略更新之间的不匹配,我们将此称为噪声事实信用分配,并将其分解为信用定位歧义与信用可靠性歧义两个方面。为解决这些问题,我们提出了FARCA(Fact-Aligned Reliability-Aware Credit Assignment,即事实对齐可靠性感知信用分配),这是一种将事实监督转化为局部化、可靠性加权的 token 级训练信号的策略优化框架。FARCA通过将事实验证的粒度与策略更新的粒度对齐,实现了细粒度的信用定位;它还引入了反事实证据归因,利用事实判断对关键证据的依赖性作为验证可靠性的经验代理,以计算可靠性权重,这些权重会调节事实奖励与局部策略优势,从而降低潜在不可靠信号对策略优化的影响。在不同模型和多个事实推理基准上的实验表明,FARCA在显著提升模型事实性的同时,保留了通用推理能力。
英文摘要:
To reduce the hallucination risk caused by outcome-driven rewards in large language models trained through reinforcement learning with verifiable rewards, existing mitigation approaches introduce process-level factual supervision. However, due to coarse-grained aggregation of factual signals and the lack of reliability assessment for these signals, they create a mismatch between fact verification and policy updates. We term this noisy factual credit assignment and decompose it into two aspects: credit localization ambiguity and credit reliability ambiguity. To address these issues, we propose FARCA (Fact-Aligned Reliability-Aware Credit Assignment), a policy optimization framework that transforms factual supervision into localized, reliability-weighted token-level training signals. FARCA achieves fine-grained credit localization by aligning the granularity of fact verification with that of policy updates. It further introduces counterfactual evidence attribution, which uses the dependence of a factual judgment on key evidence as an empirical proxy for verification reliability to compute reliability weights. These weights modulate factual rewards and local policy advantages, reducing the influence of potentially unreliable signals on policy optimization. Experiments across different models and multiple factual reasoning benchmarks show that FARCA significantly improves model factuality while preserving general reasoning capabilities.