arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ElementCheck:基于句子元素的复杂度感知长文本事实性评估

ElementCheck: Complexity-Aware Long-Form Text Factuality Evaluation via Sentence Elements

Xinming Wang, Haoran Du, Yi Chen, Jian Xu, Hongming Yang, Han Hu, Yulong Chen, Cheng-Lin Liu, Xu-Yao Zhang

arXiv 2608.26118首次发表:更新:

发表机构

Institute of Automation, Chinese Academy of Sciences; Zhongguancun Academy; Fudan University; Tencent; University of Cambridge; University of Aberdeen(中国科学院自动化研究所; 中关村学院; 复旦大学; 腾讯; 剑桥大学; 阿伯丁大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ElementCheck是一种基于句子元素的复杂度感知框架,通过提取实体对构建元素图,在FastFact-Sent等基准上提升了五个主干模型的长文本事实性验证性能,且优化了准确率与成本的权衡。

AI 中文摘要

现有的长文本事实性评估依赖分解-检索-验证流程,但该流程存在主张分解产生的噪声以及固定验证粒度的问题,导致结果不可靠。我们提出ElementCheck,这是一种基于句子元素验证长文本输出的复杂度感知框架。ElementCheck并非将句子统一分解为原子子主张,而是提取原句中通过可验证连接显式关联的实体对作为元素,并将其组织为元素图。图拓扑结构提供了用于估计句子复杂度的结构信号,支持对简单句子直接验证,对复杂句子则进行针对性的元素级细化与验证。为支持细粒度评估,我们将FastFact-Bench中的孤立主张映射回其源句子,构建了新基准FastFact-Sent。在FastFact-Sent及两个领域特定基准上的实验表明,ElementCheck在五个主干模型上均提升了事实性验证性能,同时保持了良好的准确率-成本权衡。进一步分析显示,复杂度感知验证减少了不必要的重新验证,并在不同主干模型间保持了稳定性。

英文摘要

Existing long-form factuality evaluation relies on the decompose-retrieve-verify pipeline. However, the pipeline suffers from noise from claim decomposition and fixed verification granularity, resulting in unreliable results. We propose ElementCheck, a complexity-aware framework that verifies long-form outputs via sentence elements. Instead of uniformly decomposing sentences into atomic sub-claims, ElementCheck extracts entity pairs that are explicitly linked through verifiable connections in the original sentence as elements, and organizes these into an element graph. The graph topology provides a structural signal for estimating sentence complexity, enabling direct verification for simple sentences and targeted element-level refinement and verification for complex ones. To support fine-grained evaluation, we construct a new benchmark \textbf{FastFact-Sent} by mapping isolated claims from FastFact-Bench back to their source sentences. Experiments on FastFact-Sent and two domain-specific benchmarks show ElementCheck consistently improves factuality verification across five backbone models while maintaining a favorable accuracy-cost trade-off. Further analyses demonstrate that complexity-aware verification reduces unnecessary re-verification and maintains stability across different backbones. The code is available at \href{https://github.com/gudehhh666/elementcheck.git}{Here}.

CommentsEMNLP2026 Findings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑