从判决到过程:面向多阶段事实核查的智能体强化学习
From Verdict to Process: Agentic Reinforcement Learning for Multi-Stage Fact Verification
- School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出ProFact框架,通过智能体强化学习端到端优化多阶段事实核查流程,引入过程感知奖励解决稀疏延迟监督问题,提升验证性能和推理效率。
AI中文摘要:
最近结合大型语言模型(LLMs)与检索增强推理的方法在自动化事实核查中显示出前景。为了处理复杂声明,这些核查流程通常执行多阶段工作流,协调紧密耦合的模块,包括声明分解、证据收集和判决预测。然而,现有方法孤立地优化各个阶段或依赖固定启发式规则,这限制了阶段间的自适应协调,并可能导致次优结果。在这项工作中,我们提出ProFact,一种用于多阶段事实核查轨迹端到端优化的智能体强化学习框架。ProFact训练一个统一策略来协调声明分解、证据寻找、答案生成和判决预测。为了解决最终真实性标签提供的稀疏且延迟的监督,ProFact引入了过程感知奖励,在整个核查过程中提供阶段级学习信号。实证评估表明,ProFact在验证性能和推理效率上均持续优于强基线。这些结果凸显了过程感知轨迹优化对多阶段事实核查的有效性。
英文摘要:
Recent approaches combining Large Language Models (LLMs) with retrieval-augmented reasoning have shown promise for automated fact verification. To process complex claims, these verification pipelines typically execute multi-stage workflows that coordinate tightly coupled modules, including claim decomposition, evidence gathering, and verdict prediction. However, existing methods optimize individual stages in isolation or rely on fixed heuristics, which limits adaptive coordination among stages and can lead to suboptimal outcomes. In this work, we propose ProFact, an agentic reinforcement learning framework for end-to-end optimization of multi-stage fact verification trajectories. ProFact trains a unified policy to coordinate claim decomposition, evidence seeking, answer generation, and verdict prediction. To address the sparse and delayed supervision provided by final veracity labels, ProFact introduces process-aware rewards that provide stage-level learning signals throughout the verification process. Empirical evaluation shows that ProFact consistently outperforms strong baselines in both verification performance and inference efficiency. These results highlight the effectiveness of process-aware trajectory optimization for multi-stage fact verification.