arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种用于检测政治新闻中时间不一致性的可解释一致性得分

An Explainable Coherence Score for Detecting Temporal Inconsistencies in Political News

Marius Nicusor Pantea, Adrian Groza

arXiv 2608.29175首次发表:更新:

发表机构

Artificial Intelligence Research Institute; Technical University of Cluj-Napoca(人工智能研究所; 克卢日-纳波卡技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出时间一致性得分(TCS),通过四阶段流程检测政治新闻中的时间不一致性,在100篇含时间错误的新闻基准上精确率达0.909,为事实核查提供可解释辅助。

AI 中文摘要

时间不一致性(如将指令归因于其实际区间之外、事件在发生前被表述为已发生、因果顺序颠倒等)是一种逃避基于风格的假新闻检测器的政治虚假信息:一篇措辞精美的文章即使只有一个错误日期,也没有虚假的词汇信号。本文提出了时间一致性得分(TCS),这是一种连续的、具有内在可解释性的指标,用于量化新闻文章的时间一致性,通过四阶段流程计算:提取时间事实、构建时间知识图谱、针对内部一致性规则和外部参考源进行分层验证,以及结合自动生成的解释进行得分聚合。验证过程结合了从Allen区间代数衍生的8个内部检查器,以及一个五级外部层级,范围从包含1256条精选政治事实的本地存储参考知识库,到实时Wikidata SPARQL查询。在包含注入时间错误的100篇政治新闻文章的基准测试中,该系统在选定的操作阈值下达到了0.909的精确率,仅有1个残留的误报,这一配置是专门为人机协作事实核查辅助而调整的,其中误报比漏检的成本更高。与仅输出二元标签的词汇基线不同,每一篇被标记的文章都附带不一致类型、涉及的实体以及与该主张相矛盾的参考源。

英文摘要

Temporal inconsistencies, such as mandates attributed outside their real interval, events presented as past before they occurred, or inverted causal sequences, are a form of political disinformation that evades style-based fake news detectors: a well-written article with a single wrong date carries no lexical signal of falsehood. This paper introduces the Temporal Coherence Score (TCS), a continuous, intrinsically interpretable metric that quantifies the temporal coherence of a news article, computed by a four-stage pipeline: extraction of temporal facts, construction of a temporal knowledge graph, hierarchical verification against internal consistency rules and external reference sources, and score aggregation with automatically generated explanations. Verification combines eight internal checkers derived from Allen's interval algebra with a five-level external hierarchy ranging from a locally stored reference knowledge base of 1{,}256 curated political facts to live Wikidata SPARQL queries. On a benchmark of 100 political news articles with injected temporal errors, the system reaches a precision of 0.909 at the selected operating threshold, with a single residual false positive, a profile deliberately tuned for human-in-the-loop fact-checking assistance, where false alarms are costlier than missed detections. Unlike lexical baselines that output only a binary label, every flagged article is accompanied by the inconsistency type, the entities involved, and the reference source that contradicts the claim.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑