超越裁决:科学事实核查中人类与大语言模型推理的基于图的分析
Beyond Verdicts: A Graph-Based Analysis of Human and LLM Reasoning in Scientific Fact-Checking
- Interdisciplinary Transformation University Austria (ITU)(奥地利跨学科转型大学(ITU))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出类型化推理图框架,用于对比科学事实核查中人类与LLM的推理路径,基于MISSCIPLUS数据集评估GPT-5等模型,发现各模型在裁决性能与人类对齐度上存在差异。
AI中文摘要:
引用合法论文的错误信息在歪曲这些研究实际报告内容时危害尤其大。尽管现有的基于大语言模型(LLM)的自动事实核查系统能够评估模型是否给出错误裁决并生成该决策的解释,但它们通常无法表明模型是否遵循与人类专家相同的推理路径,还是通过不同但仍有效的路径得出裁决。在本研究中,我们引入了一种基于图的框架(类型化推理图),用于比较科学事实核查中人类与LLM的推理路径。基于生物医学错误信息中错误推理的现有研究MISSCIPLUS(Glockner等人,2025),我们将每个解释建模为一个推理图,该图将错误主张与相关研究背景、研究发现、支持谬误的前提以及谬误标签关联起来。这种表示方式能够在特定谬误子图的层面上实现人类与LLM推理的一一对应。对于与人类推理不对齐的LLM路径,我们验证其是否锚定在引用的研究中、与主张相关且对裁决具有充分性。使用来自MISSCIPLUS的84个错误主张,我们在不同提示和证据设置下评估了GPT-5、Claude Opus 4.7和Qwen3-32B。结果显示出不同的性能维度:Qwen3-32B的裁决失败率最低,GPT-5的人类对齐度最高,而Claude Opus 4.7的裁决预测能力较弱,但在成功案例中往往具有有效的推理。
英文摘要:
Misinformation that cites legitimate papers can be especially harmful when it distorts what those studies actually report. While existing automatic fact-checking systems based on large language models (LLMs) can assess whether a model assigns an Incorrect verdict and can gen- erate explanations for that decision, they typi- cally do not indicate whether the model follows the same reasoning path as human experts or arrives at the verdict through a different but still valid path. In this work, we introduce a graph- based framework (typed reasoning graph) for comparing human and LLM reasoning paths in scientific fact-checking. Building on prior work on fallacious reasoning in biomedical misinformation, MISSCIPLUS (Glockner et al., 2025), we model each explanation as a rea- soning graph that links the false claim to the relevant study context, study findings, fallacy- supporting premises, and fallacy labels. This representation enables one-to-one alignment of human and LLM reasoning at the level of fallacy-specific sub-graphs. For non-human- aligned LLM paths, we validate grounding in the cited study, relevance to the claim, and suf- ficiency for the verdict. Using 84 false claims from MISSCIPLUS, we evaluate GPT-5, Claude Opus 4.7, and Qwen3-32B across prompt and evidence settings. Results show distinct perfor- mance dimensions: Qwen3-32B has the lowest verdict failure rate, GPT-5 the highest human alignment, and Claude Opus 4.7 weak verdict prediction but often valid reasoning in success- ful cases