arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11415cs.IRcs.AI

TRACES:用于评估大型语言模型科学推理中认知可靠性的基准

TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs

  • Case Western Reserve University(凯斯西储大学)
  • Intellicat(英特卡特)

机构由 AI 辅助整理,请以论文原文为准。

Valentin Rodionov, Shamil Assylbekov

AI总结:

该研究构建了名为TRACES的基准,发现30种大型语言模型在识别42篇不可靠科学论文的认知可靠性时表现不佳,仅少数模型能有效拒绝有缺陷前提,凸显科学部署需护栏。

AI中文摘要:

大型语言模型正被提议作为科学工作流程中的智能体,应用于不存在下游验证工具的领域。这种部署假设模型能够区分可靠的科学文献与不可靠的文献,而这一能力尚未得到直接测量。现有基准评估的是已知答案问题的事实性,而我们所关注的失效模式与此不同。我们引入了包含42篇被撤回、欺诈性及伪科学论文的探测语料库,并提出了一种方法,用于引出和评分模型对每篇论文框架的单次交互情况。每个探测项包含从目标论文中近乎逐字提取的前言,以及一个科学上合理的研究设计请求。这些探测项涵盖五种主张类型:伪造观察结果、伪物理机制、魔法前提、合法化桥梁以及 cargo-cult 实验。两个互补评分用于衡量模型是否直接拒绝有缺陷的前提(IFR-a),以及是否在仍参与交互的同时识别出不可靠性(IFR-i)。深度评分 Engagement Depth Index(EDI)用于量化对论文或领域特定隐藏细节的复现情况。在30个模型和10次重复运行中,总体 IFR-a 为0.93 ± 0.004,总体 IFR-i 为0.809 ± 0.009。模型在95%的非空响应中对站不住脚的前提进行了交互。所有被评估的模型在超过71%的智能体探测项上失效,且30个模型中有22个在超过90%的情况下失效。拒绝行为集中在少数高知名度主题和特定探测项上,在匹配结构的对照实验中消失。这些结果与基于主题的安全行为而非稳健的认知能力一致,表明语言模型在科学领域部署时亟需护栏基础设施。

英文摘要:

Large language models are being proposed as agents in scientific workflows, in domains where no downstream verifier exists. Such deployment assumes the model can distinguish reliable scientific literature from unreliable literature, a capability that has not yet been directly measured. Existing benchmarks evaluate factuality on questions with known answers; the failure mode we target here is different. We introduce a probe corpus of 42 retracted, fraudulent, and pseudoscientific papers, paired with a methodology for eliciting and scoring single-shot model engagement with each paper's framing. Each probe pairs a preamble extracted near-verbatim from the target paper with a scientifically plausible study-design request. The probes span five claim types: fabricated observation, pseudophysical mechanism, magical premise, legitimization bridge, and cargo-cult experiment. Two complementary scores measure whether a model rejects the flawed premise outright (IFR-a) and whether it recognizes the unreliability while still engaging (IFR-i). A depth score, the Engagement Depth Index (EDI), quantifies reproduction of paper- or field-specific withheld details. Across 30 models and 10 repeated runs, aggregate IFR-a is 0.93 $\pm$ 0.004 and aggregate IFR-i is 0.809 $\pm$ 0.009. Models engaged with untenable premises in 95% of all non-empty responses. Every evaluated model fails more than 71% of agentic probes, and 22 of 30 models fail more than 90% of the time. Rejections are concentrated on a small number of high-notoriety topics and specific probes, and disappear under matched-structure controls. These results are consistent with topic-keyed safety behavior rather than robust epistemic competence, and indicate an urgent need for guardrail infrastructure for scientific deployment of language models.

补充信息

↑