arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.20527cs.AI

评估和保障智能科学合成中的引用忠实性

Evaluating and Guarding Citation Faithfulness in Agentic Scientific Synthesis

Taewan Goo, Junsik Kim, Kyulhee Han, GwonYul Jo, Jong-Soo Kim, Tae-Hyung Kim

首次发表
浏览论文内容

中文总结 AI 辅助

研究智能语言模型系统引用答案检查的可靠性问题,提出黄金锚定评估协议和可部署防护措施,能验证验证者、测量重新归因,为无支撑引用设限,经多模型和管道验证后作为单GPU工具包发布。

中文摘要 AI 辅助

诸如OpenScholar和PaperQA2等智能语言模型系统阅读科学文献并返回引用答案,它们及其基准测试已经使用固定归因模型或人工评分来检查这些引用是否成立,但都未审核该检查本身的可靠性。研究表明这种检查不可靠且很重要。相同智能体输出的无支撑引用率仅取决于验证者的严格程度,在约3%至约18%之间,且验证者对哪些引用得到支持意见一致,但对标记哪些引用存在分歧。为此提出了黄金锚定评估协议和可部署防护措施,协议能验证验证者、测量重新归因并校准针对人类黄金标准而非其他模型裁决的保证;防护措施添加了一个分裂共形层,对未被选定标记规则检测到的真正无支撑引用设置无分布、有限样本界限。在四个公开的27 - 35B模型和三个智能管道上进行了验证,并作为一个开放的单GPU工具包发布。

英文摘要

Agentic LLM systems such as OpenScholar and PaperQA2 read the scientific literature and return cited answers, and both they and their benchmarks already check whether those citations hold, with a fixed attribution model or human graders. Neither audits the reliability of that check itself. We show it is not reliable, and that this matters. On identical agent outputs the measured unsupported-citation rate ranges from about 3% to about 18% depending only on the verifier's strictness, and although verifiers agree on which citations are supported, they disagree on which to flag (negative-specific agreement 0.27 to 0.30), so no single flag set is trustworthy and cross-paper comparison is invalid without a named verifier and protocol. We present a gold-anchored evaluation protocol and a deployable guard that make this behavior measurable and bounded. The protocol validates the verifier, measures re-attribution, and calibrates a guarantee against human gold rather than another model's verdict; the verifier is a swappable instrument chosen on cost (recall 0.94 on the supported class, held out), and re-attribution is a commodity step where a deterministic BM25 matches the best open generator. The guard adds a split-conformal layer placing a distribution-free, finite-sample bound on truly unsupported citations that slip past a chosen flagging rule, a guarantee on catch rate rather than conclusion correctness. The bound holds on held-out gold, and we identify and quantify the condition governing its transfer to deployment, calibration-negative difficulty, with a concrete recalibration recipe, left untested by prior conformal-factuality work. Validated across four open 27-35B models and three agentic pipelines on public benchmarks (SciFact, QASA, PubMedQA), with confidence intervals on every headline number, the protocol and guard ship as an open single-GPU kit.

发表机构

  • Seoul National University(首尔国立大学)
  • BioNexus(生物 Nexus)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑