AI 中文总结
OpenDiscoveryTrace提供558条AI科学智能体轨迹,通过9字段过程追踪揭示仅输出评估无法发现的行为差异,支持过程级评估与智能体审计。
AI 中文摘要
现有的自主AI科学家基准仅评估最终输出——生成的代码、假设或论文——却丢弃了获得这些输出的推理过程。这使得无法审计科学方法论、诊断失败模式,或区分系统性推理与侥幸猜测。我们提出OpenDiscoveryTrace,一个包含558条完整AI科学智能体轨迹的公开数据集,它捕捉模型如何推理,而不仅仅是它们产生了什么。每条轨迹记录了一个结构化的每步9字段的跟踪信息——包括思考、工具调用、观察、错误、修订触发器和自我报告的置信度——因为模型执行了124项科学任务,涵盖药物发现、材料科学、基因组学和科学文献分析。该数据集覆盖七个模型:三个前沿模型(GPT-5.4、Claude Opus 4.6和Gemini 3.1 Pro;各124条轨迹,在领域和难度级别上完全平衡)和四个开放权重模型(Qwen2.5-7B、Mistral-7B-v0.3、Phi-3.5-mini和Qwen2.5-1.5B;各30条),外加60条实时检索变体轨迹。对363条LLM评判轨迹的初步分析显示,过程轨迹暴露了仅输出评估无法看到的行为差异:所有三个前沿模型都取得了相当的成功率(84-89%),但Claude Opus 4.6产生的错误比GPT-5.4多30倍(每条轨迹2.5个对0.08个,p < 0.0001,Cliff's δ = 0.613),且错误概况定性不同——Claude有66.7%的工具误用,而GPT-5.4有83.6%的推理错误。我们定义了五个基准任务,使用逻辑回归、随机森林、LSTM和Transformer模型的基线。该数据集、轨迹模式、智能体框架和基准定义在CC BY 4.0下公开可用,以支持过程级评估、科学智能体审计和AI治理方面的研究。
英文摘要
Existing benchmarks for autonomous AI scientists evaluate only final outputs---generated code, hypotheses, or papers---yet discard the reasoning process by which those outputs were obtained. This makes it impossible to audit scientific methodology, diagnose failure modes, or distinguish systematic reasoning from fortunate guessing. We present \textbf{OpenDiscoveryTrace}, a public dataset of 558 complete AI scientific agent trajectories that captures how models reason, not just what they produce. Each trajectory records a structured 9-field-per-step trace---including thoughts, tool calls, observations, errors, revision triggers, and self-reported confidence---as models execute 124 scientific tasks spanning drug discovery, materials science, genomics, and scientific literature analysis. The dataset covers seven models: three frontier models (GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro; 124 trajectories each, fully balanced across domains and difficulty levels) and four open-weight models (Qwen2.5-7B, Mistral-7B-v0.3, Phi-3.5-mini, and Qwen2.5-1.5B; 30 each), plus 60 live-retrieval variant trajectories. Pilot analysis on 363 LLM-judged trajectories reveals that process traces expose behavioral differences invisible to output-only evaluation: all three frontier models achieve comparable success rates (84--89%), yet Claude Opus 4.6 produces 30$\times$ more errors than GPT-5.4 (2.5 vs. 0.08 per trajectory, $p < 0.0001$, Cliff's $δ= 0.613$), with qualitatively different error profiles---66.7% tool misuse for Claude versus 83.6% reasoning errors for GPT-5.4. We define five benchmark tasks with baselines from logistic regression, random forests, LSTMs, and Transformer models. The dataset, trace schema, agent harness, and benchmark definitions are publicly available under CC BY 4.0 to support research on process-level evaluation, scientific agent auditing, and AI governance.
CommentsBest Dataset Award, ICML 2026 Workshop on AI for Science (AI Scientists: Tools, Co-authors, or Founders?). Code: https://github.com/aayambansal/OpenDiscoveryTrace Dataset: https://huggingface.co/datasets/aayambansall/OpenDiscoveryTrace