发表机构
Carnegie Mellon University; Syracuse University(卡内基梅隆大学; 雪城大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出证据账本裁决的主张-证据可追溯性工作流,构建2335行盲基准实验,结果显示智能体证据账本条件的关系准确率和宏F1显著优于非智能体基线,可退回需处理的主张并为AI辅助写作提供可审计可追溯层。
AI 中文摘要
智能体生成主张的速度快于作者核查引用或检索证据是否支持这些主张的速度。我们研究证据账本裁决:一种主张-证据可追溯性工作流,将每项主张与证据包配对,分配支持关系,并将未获支持、与证据矛盾或混合证据的主张退回作者。实证核心是一个2335行的盲基准,由AVeriTeC、CLIMATE-FEVER和SciFact中的独立外部标签构建。预测期间隐藏真实关系和源证据标签,仅在评分时结合。在该基准上,智能体证据账本条件达到0.676的关系准确率和0.601的宏F1值,而最佳非智能体基线的准确率为0.383、宏F1为0.303。它还将1270/1435项真实标签显示为矛盾、证据缺失或混合证据的主张退回,同时退回295/900项获支持的主张。这些结果表明,证据账本裁决可将异构证据包转化为AI辅助写作的可审计可追溯层。
英文摘要
AI agents can draft claims faster than authors can check whether the cited or retrieved evidence supports them. We study evidence-ledger adjudication: a claim-evidence traceability workflow that pairs each claim with an evidence packet, assigns a support relation, and routes unsupported, contradicted, or mixed-evidence claims back to the author. The empirical core is a 2,335-row blind benchmark built from independent external labels in AVeriTeC, CLIMATE-FEVER, and SciFact. Gold relations and source evidence labels are hidden during prediction and joined only for scoring. On this benchmark, the agent evidence-ledger condition achieves 0.676 relation accuracy and 0.601 macro-F1, compared with 0.383 accuracy and 0.303 macro-F1 for the best non-agent baseline. It also routes 1270/1435 claims whose gold labels indicate contradiction, missing evidence, or mixed evidence, while routing 295/900 supported claims. These results show that evidence-ledger adjudication can turn heterogeneous evidence packets into an auditable traceability layer for AI-assisted writing.