arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22223cs.CLcs.LG

EAVer: 长文本事实性验证作为端到端智能体策略

EAVer: Long-Form Factuality Verification as an End-to-End Agentic Policy

  • University of Illinois Chicago(伊利诺伊大学芝加哥分校)
  • Taobao & Tmall Group(淘宝天猫集团)
  • Shandong University(山东大学)
  • Peking University(北京大学)
  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

Kening Zheng, Aoying Zheng, Zhigang Chang, Yazhi Guo, Miaotian Guo, Qingwei Zong, Xianhai Xie, Weiqiang Jin, Chengze Li, Hanrong Zhang, Jie Yang, Wei-Chieh Huan… 展开作者

Kening Zheng, Aoying Zheng, Zhigang Chang, Yazhi Guo, Miaotian Guo, Qingwei Zong, Xianhai Xie, Weiqiang Jin, Chengze Li, Hanrong Zhang, Jie Yang, Wei-Chieh Huang, Lingzhe Zhang, Liancheng Fang, Xin Zou, Hanqian Li, Jiahao Huo, Yibo Yan, Zizhuang Deng, Lei Miao, Wei Guo, Haihong Tang, Bo Zheng, Philip S. Yu

AI总结:

EAVer提出端到端智能体策略,统一管理长文本事实性验证流程,通过声明分组、置信度路由和上下文证据复用,显著减少搜索次数并提升验证性能,且跨模型规模泛化良好。

AI中文摘要:

长文本事实性验证通常被实现为静态的分解-搜索-验证流水线,其中独立提示的模块处理声明并调用外部搜索。将声明独立处理使得大语言模型和搜索调用次数随声明数量线性增长,并导致对相关声明的重叠证据进行重复搜索。我们引入了EAVer,一种端到端智能体验证器,它学习将完整的响应级验证工作流控制为一个统一策略。EAVer将语义相关的声明分组,根据置信度将每组路由到直接验证或定向搜索,并将搜索返回的证据保存在紧凑的上下文备忘录中,以供跨声明重用。为了训练该策略,我们开发了一个特权教师合成流水线,该流水线将黄金声明注释转换为可执行的多轮工具交互轨迹,使用实时搜索而非事后理由。结构、标签对齐、工具使用、搜索预算和泄漏检查产生了1,447条质量受控的轨迹。我们进一步构建了794对双向同轨迹偏好对,保持声明分组、搜索和证据固定,从而能够在事实性决策令牌上进行决策聚焦的直接偏好优化(DPO)。使用Qwen3-8B的结果表明,EAVer在每个基准上均优于最强的基于搜索的基线,在VeriFastScore上高出2.88个Macro-F1点,在分布外FaStFact-Bench上高出4.73个点,同时比搜索效率最高的基线少使用约80%的搜索次数。此外,EAVer在4B到32B参数的各种模型上持续提升性能,展示了其强大的泛化能力。

英文摘要:

Long-form factuality verification is commonly implemented as a static decompose-search-verify pipeline, with separately prompted modules processing claims and invoking external search. Treating claims independently makes LLM and search calls scale with claim count and causes repeated searches for overlapping evidence about related claims. We introduce EAVer, an End-to-end Agentic Verifier that learns to control the complete response-level verification workflow as a unified policy. EAVer groups semantically related claims, routes each group to direct verification or targeted search based on confidence, and keeps evidence returned by search in compact in-context memos for cross-claim reuse. To train this policy, we develop a privileged-teacher synthesis pipeline that converts gold claim annotations into executable multi-turn tool-interaction trajectories with live search rather than post-hoc rationales. Structural, label-alignment, tool-use, search-budget, and leakage checks yield 1,447 quality-controlled trajectories. We further construct 794 bidirectional same-trajectory preference pairs that keep claim grouping, search, and evidence fixed, enabling decision-focused Direct Preference Optimization (DPO) over factuality-decision tokens. The results with Qwen3-8B show that EAVer outperforms the strongest search-based baseline on each benchmark by 2.88 Macro-F1 points on VeriFastScore and 4.73 points on the out-of-distribution FaStFact-Bench, while using about 80% fewer searches than the most search-efficient baseline. Moreover, EAVer consistently improves performance across models ranging from 4B to 32B parameters, demonstrating its strong generalizability.

↑