AI 中文总结
TruthInsightBench是为科学发现设计的基准,其评估显示现有编码智能体的瓶颈是科学判断而非编码,真正的科学发现仍未实现。
AI 中文摘要
人们越来越多地将自主编码智能体提议为开展分析并撰写研究报告的AI科学家系统,但执行规定的分析与做出发现并非同一回事。现有的基准是为复现而配置的:任务、数据和评分规则围绕隐藏的目标研究构建,奖励的是对其结果的复现。我们提出TruthInsightBench,一个为发现而配置的基准。它的40个盲任务来自10个科学领域的40项同行评审研究,仅公开中性的科学目标和固定数据;源结论、预期值和分析路径被隐藏,让智能体自行确定数据支持何种主张。一个基于固定大语言模型(LLM)的评判者从6个维度对智能体自身主张的证据成熟度进行评分,这些维度被操作化为29项基于人工制品的条目,采用自动化、确定性的聚合方式,且无需针对每个实例进行人工评分,因此随着智能体的发展,评估可自动重复进行。在一个固定的基础模型上,四个编码智能体形成了一个狭窄的平台(100分制下为58.4-60.3分),且在统计上没有可靠的成对区分:它们能够胜任地执行并记录分析,具备相对较强的证据可审计性和新颖性,但在很大程度上缺乏确立可信主张所需的辨别性行为(控制、稳健性、可证伪性及跨数据集泛化能力)。瓶颈在于科学判断而非编码,真正的发现仍遥不可及。TruthInsightBench将这一差距转化为可衡量的目标;数据和评分代码位于此https URL。
英文摘要
Autonomous coding agents are increasingly proposed as AI-scientist systems that conduct analyses and write research reports, but executing a prescribed analysis is not the same as making a discovery. Existing benchmarks are configured for reproduction: tasks, data, and rubrics are built around a hidden target study, and recovery of its result is rewarded. We present TruthInsightBench, a benchmark configured for discovery. Its 40 blind tasks, drawn from 40 peer-reviewed studies across 10 scientific domains, expose only a neutral scientific objective and frozen data; source conclusions, expected values, and analysis paths are withheld, leaving the agent to determine what claim the data support. A fixed LLM-based judge scores the evidentiary maturity of an agent's own claims along six dimensions, operationalized as 29 artifact-grounded items, with automated, deterministic aggregation and no per-instance human grading, so evaluation can be repeated automatically as agents evolve. On one frozen base model, four coding agents form a narrow plateau (58.4-60.3 of 100) with no statistically reliable pairwise separation: they execute and document analyses competently, with comparatively strong evidence auditability and novelty, but largely lack the discriminating acts that establish a trustworthy claim (controls, robustness, falsifiability, and cross-dataset generalization). The bottleneck is scientific judgment rather than coding, and genuine discovery remains out of reach. TruthInsightBench makes this gap a measurable target; data and scoring code are at https://github.com/TruthInsight-stack/TruthInsightBench.
Comments27 pages, 7 tables, 5 figures