arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

验证前的理解:用于自动引文验证的声明规范化

Understanding before verifying: Claim normalization for automated citation verification

Yifan He, Mengjia Wu, Siming Deng, Yi Zhang

arXiv 2608.30145首次发表:更新:

AI 中文总结

针对现有两阶段引文验证框架的缺陷,提出声明规范化方法,开发三阶段框架CNCV,经实验使编码器和生成式LLM的宏F1分别平均提升12%和10%。

AI 中文摘要

引文准确性因对研究可靠性的重要性已被研究数十年,内容级引文验证用于评估学术声明的可靠性。近期研究采用继承自事实核查的两阶段检索-分类框架,但该设计忽略了原始引用声明的复杂性,为验证系统引入了范围不匹配、视角不匹配与命题纠缠三个问题,这些问题增加了检索与分类的难度,从而限制了模型性能。针对这一缺口,我们提出声明规范化,在检索与分类前对原始引用声明应用三种改写策略,使每个下游模型能执行单一、定义明确的任务。基于此方法,我们开发了新的三阶段框架CNCV(声明规范化引文验证),由声明规范化、带 grounding 的证据检索及引文分类组成。我们在人工标注的引文实例上进行析因实验,用18种分类器评估CNCV。与现有两阶段框架相比,CNCV使编码器的宏F1平均提升12%,生成式大语言模型(LLM)的宏F1平均提升10%,这源于实验中确定的主导因素——证据质量的提升;从自动规范化声明中检索的证据,其下游分类性能与使用人工标注证据获得的性能在统计上等效。

英文摘要

Citation accuracy has been studied for decades because of its importance to research reliability. Content-level citation verification assesses the reliability of scholarly claims. Recent work adopts a two-stage retrieval-classification framework inherited from fact-checking. However, this design overlooks the complexity of the raw citing claim and introduces three issues into the verification system, namely scope mismatch, perspective mismatch, and proposition entanglement. These issues increase the difficulty of retrieval and classification, thereby limiting model performance. Motivated by this gap, we propose claim normalization, which applies three rewriting strategies to the raw citing claim before retrieval and classification, allowing each downstream model to perform a single, well-defined task. Building on this method, we develop Claim-Normalized Citation Verification (CNCV), a new three-stage framework consisting of claim normalization, evidence retrieval with grounding, and citation classification. We evaluate CNCV across 18 classifiers using a factorial experiment on human-annotated citation instances. Compared with the prior two-stage framework, CNCV improves macro F1 by an average of 12% for encoders and 10% for generative LLMs, driven by improved evidence quality, the dominant factor identified in our experiments. Evidence retrieved from automatically normalized claims yields downstream classification performance statistically equivalent to that obtained with manually annotated evidence.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑