arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05302cs.IR

用于提升AI生成新闻摘要事实可靠性的跨平台认知验证

Cross-platform epistemic verification for improving factual reliability in AI-generated news summarization

Zhuo Xie, Haoze Ni

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出多源证据共识验证(MECV)框架,通过多异构来源证据聚合与多LLM陪审团机制修正AI生成新闻摘要的幻觉,在SummEdits基准上提升了事实一致性,为可信AI与自动化新闻研究提供了新方法。

中文摘要 AI 辅助

本研究提出了多源证据共识验证(Multi-source Evidence Consensus Verification, MECV),这是一种用于AI生成新闻摘要的事后幻觉修正框架。MECV不依赖单一检索通道,而是聚合来自多个异构来源的证据,包括源文档、维基百科和开放网络检索。该框架还引入了多LLM陪审团机制,通过验证模型间的矛盾感知共识评分来评估事实可靠性。被识别为可能缺乏支持的主张会通过迭代式最小编辑优化进行修正。该框架在SummEdits基准上进行评估,使用GPT-4o-mini和DeepSeek-Chat作为验证陪审团,Qwen-Plus作为协调器。实验结果表明,MECV在保留原始摘要语义结构的同时提升了事实一致性。研究结果进一步表明,异构证据来源间的一致性可作为识别AI生成摘要中事实不确定性的有用信号,包括在金融新闻聚合等信息敏感领域。本研究通过引入用于幻觉修正的多源验证框架,为可信AI和自动化新闻研究做出了贡献,并证明了基于共识的验证对提升AI生成新闻摘要事实可靠性的价值。

英文摘要

This study proposes Multi-source Evidence Consen- sus Verification (MECV), a post-hoc hallucination cor- rection framework for AI-generated news summariza- tion. Instead of depending on a single retrieval channel, MECV aggregates evidence from multiple heterogeneous sources, including the source document, Wikipedia, and open-web retrieval. The framework further incorporates a multi-LLM jury mechanism that estimates factual reliabil- ity through contradiction-aware consensus scoring across verifier models. Claims identified as potentially unsup- ported are revised through iterative minimal-edit refine- ment. The proposed framework is evaluated on the SummEd- its benchmark using GPT-4o-mini and DeepSeek-Chat as the verifier jury, with Qwen-Plus as the orchestra- tor. Experimental results show that MECV improves fac- tual consistency while preserving the semantic structure of the original summaries. The findings further suggest that agreement across heterogeneous evidence sources can serve as a useful signal for identifying factual uncertainty in AI-generated summaries, including in information- sensitive domains such as financial news aggregation. This study contributes to research on trustworthy AI and automated journalism by introducing a multi-source verification framework for hallucination correction and demonstrating the value of consensus-based verification for improving factual reliability in AI-generated news summarization.

↑