arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当事实验证分数提升时会发生什么变化?训练验证器与LLM的证据及答案核算

What Changes When Fact-Verification Scores Improve? Evidence and Answer Accounting Across Trained Verifiers and LLMs

Han Chen, Yingrui Li

arXiv 2609.27064首次发表:更新:

AI 中文总结

本研究通过联合事实验证分数提升的分解分析,量化了证据替换对分数增益的贡献,并考察了上下文长度和模型选择的影响,揭示了聚合指标遗漏的声明层面变化。

AI 中文摘要

联合事实验证分数同时评估答案和提交的证据。当分数提升时,若保持答案不变,其中有多少增益得以保留?在FEVEROUS上,严格分数是提交证据中答案正确且包含完整标注证据组的声明的百分比。在四个训练的DeBERTa检查点和7,890个声明上,用UnifEE证据替换DCUF证据使严格分数提高了9.61个百分点,而答案准确率仅提高了1.96个百分点。在这些检查点条件下,严格分数增益的配对95%置信区间为[8.77, 10.43]。仅替换传递给评分器的证据,在分别保留由DCUF或UnifEE证据生成的答案时,可解释7.92或9.08个百分点的增益。为了检验这一证据增益如何依赖于评估选择,我们在FEVER、FEVEROUS和SciFact上,使用两种答案格式和两种上下文预算,从两个8B LLM生成了470,400个响应。将上下文从256个token增加到2,048个token,在FEVEROUS上,Qwen和Llama的固定答案证据增益分别提高了3.84和3.10个百分点。这些效应未达到预先指定的跨数据集标准,而某些区间超出了两点小效应界限。事后分析量化了答案和提交证据的变化,并展示了聚合准确率和证据覆盖率何时遗漏了声明层面的模式。四种答案-证据分数组合揭示了端点和聚合指标无法解决的差异。

英文摘要

A joint fact-verification score assesses answers and submitted evidence together. When the score improves, how much of the gain remains if the answers are held fixed? On FEVEROUS, strict score is the percentage of claims with a correct answer and a complete annotated evidence group in the submitted evidence. Across four trained DeBERTa checkpoints and 7,890 claims, replacing DCUF evidence with UnifEE evidence raises strict score by 9.61 percentage points, compared with 1.96 percentage points in answer accuracy. The paired 95% interval for the strict-score gain is [8.77, 10.43], conditional on these checkpoints. Replacing only the evidence passed to the scorer accounts for 7.92 or 9.08 percentage points when we retain the answers generated from DCUF or UnifEE evidence, respectively. To examine how this evidence gain depends on evaluation choices, we generate 470,400 responses from two 8B LLMs on FEVER, FEVEROUS, and SciFact under two answer formats and two context budgets. Increasing context from 256 to 2,048 tokens raises the fixed-answer evidence gain on FEVEROUS by 3.84 and 3.10 percentage points for Qwen and Llama, respectively. The effects fall short of the prespecified cross-dataset criterion, while some intervals extend beyond the two-point small-effect bound. Post-hoc analyses quantify changes in answers and submitted evidence, and show when aggregate accuracy and evidence-coverage rates miss the claim-level pattern. The four answer-evidence score combinations reveal changes that endpoint and aggregate metrics leave unresolved.

Comments24 pages. Both authors contributed equally

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑