OpenFC:面向开放搜索事实核查的验证策略学习
OpenFC: Learning Verification Policies towards Open-Search Fact Checking
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
- University of Electronic Science and Technology of China(电子科技大学)
- Beijing Institute of Technology(北京理工大学)
- Nankai University(南开大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
OpenFC提出两阶段训练框架,将Qwen3-8B训练为统一验证策略控制器,在六个基准上取得最高平均准确率70.39%和宏F1 63.30%,实现开放搜索事实核查的决策优化。
AI中文摘要:
开放搜索事实核查不仅仅是检索后分类,而是一个顺序决策问题,其中每一次查询、来源访问和停止决策都会重塑可用于验证的证据。然而,现有系统通常将这些决策分散在预定义流程或单独提示的模块中,而不是将其作为统一的任务特定策略进行学习。我们提出了OpenFC,一个统一的验证策略训练框架,将Qwen3-8B后训练为紧凑的下一动作控制器,负责推理、证据获取和停止。OpenFC分两个阶段学习该策略。逐步校准冷启动(SCCS)使用强大的训练时监督器,在执行前审查初始推理、工具使用和停止提议,从而在没有金标准裁决的情况下生成可靠的轨迹用于监督微调。随后,验证感知强化学习(VA-RL)通过预算感知的工具奖励、标签感知的优势重新加权和局部响应掩码,改进冷启动策略对未解决主张的处理。在六个事实核查基准上,OpenFC实现了70.39%的平均准确率和63.30%的宏F1分数,在评估方法中总体平均值最高。分阶段消融进一步表明,SCCS和VA-RL提供了互补的增益,支持了两阶段训练框架的设计。这些结果使OpenFC成为开放搜索事实核查的强大有效框架。我们将开源代码并发布模型检查点以支持可重复性。
英文摘要:
Open-search fact checking is not merely retrieval followed by classification, but a sequential decision problem in which every query, source visit, and stopping decision reshapes the evidence available for verification. Yet existing systems often distribute these decisions across predefined pipelines or separately prompted modules rather than learning them as a unified task-specific policy. We introduce \textbf{OpenFC}, a unified verification-policy training framework that post-trains Qwen3-8B as a compact next-action controller over reasoning, evidence acquisition, and stopping. OpenFC learns this policy in two stages. \textbf{Stepwise-Calibrated Cold Start (SCCS)} uses a strong training-time supervisor to review post-initial reasoning, tool-use, and stopping proposals before execution, producing reliable trajectories for supervised fine-tuning without access to gold verdicts. \textbf{Verification-Aware Reinforcement Learning (VA-RL)} then improves the cold-start policy on unresolved claims through budget-aware tool rewards, label-aware advantage reweighting, and localized response masking. Across six fact-checking benchmarks, OpenFC achieves 70.39\% average accuracy and 63.30\% macro-F1, the highest overall averages among the evaluated methods. Stage-wise ablations further show that SCCS and VA-RL provide complementary gains, supporting the design of the two-stage training framework. These results position OpenFC as a strong and effective framework for open-search fact-checking. We will open-source our code and release the model checkpoints to support reproducibility.