arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MDArena:面向真实分子动力学工作流的编码智能体评估基准

MDArena: Evaluating Coding Agents on Realistic Molecular Dynamics Workflows

Nithishwer Mouroug Anand, Wei-Tse Hsu, Kyle Vaccaro, Eden James Gage, Jonathan David Colburn, Linda Xi Phan, Minjoon Seo, Kevin Guan, Philip C. Biggin

arXiv 2608.02642首次发表:更新:

发表机构

University of Oxford; Scripps Research Institute(牛津大学; 斯克里普斯研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MDArena基准评估了6种编码智能体在真实分子动力学任务上的表现,发现其在精细科学工作流细节上存在不足,为追踪其可靠性进展提供了平台。

AI 中文摘要

加速科学发现是AI最重要的应用之一,而计算生物分子模拟是这一努力中极具潜力的目标。编码智能体有望自动化该工作流的大量环节,但其在真实分子动力学(MD)任务上的可靠性仍未得到充分表征。为解决这一问题,我们推出MDArena,这是一个包含50个容器化任务的基准,这些任务来自活跃的生物分子模拟项目,涵盖29个分子系统和14种广泛的研究协议,包括轨迹分析、复杂系统制备、自由能协议和增强采样。我们评估了6种模型/框架配置,涵盖Codex和OpenCode。在评估的配置中,Codex GPT-5.5在超高推理 effort 下表现最佳,达到24/50的Strict-Pass@1成功率(48%),其次是Codex GPT-5.5 Medium(21/50)和OpenCode Gemini Flash 3.5(20/50)。所有配置的平均正确性和流程奖励均显著高于严格成功率,表明智能体经常取得有意义的部分进展,但在可复现科学工作流所需的精细细节上存在不足。困难任务仍基本未解决,特别是膜蛋白系统制备和炼金术自由能设置,所有评估配置均未解决或接近未解决。因此,MDArena揭示了编码智能体作为辅助助手的实用性与其作为自主MD研究人员的可靠性之间存在巨大差距,同时提供了一个可复现且可扩展的平台,用于追踪缩小这一差距的进展。

英文摘要

Accelerating scientific discovery is among the most consequential applications of AI, and computational biomolecular simulation stands out as a particularly promising target within this broader effort. Coding agents promise to automate significant portions of this workflow, yet their reliability on realistic molecular dynamics (MD) tasks remains poorly characterized. To address this issue, we introduce MDArena, a benchmark of 50 containerized tasks drawn from active biomolecular simulation projects, spanning 29 molecular systems and 14 broad research protocols, including trajectory analysis, complex system preparation, free-energy protocols, and enhanced sampling. We evaluate six model/harness configurations spanning Codex and OpenCode. Among the evaluated configurations, Codex GPT-5.5 at extra-high reasoning effort performs best, reaching 24/50 Strict-Pass@1 successes (48%), followed by Codex GPT-5.5 Medium with 21/50, and OpenCode Gemini Flash 3.5 with 20/50. Average correctness and process rewards are substantially higher than strict success rates across all configurations, indicating that agents frequently make meaningful partial progress but fail on the fine-grained details required for reproducible scientific workflows. Hard tasks remain largely unsolved, particularly membrane-protein system preparation and alchemical free-energy setup, both unsolved or near-unsolved by every evaluated configuration. MDArena thus exposes a substantial gap between the usefulness of coding agents as supervised assistants and their reliability as autonomous MD researchers, while providing a reproducible and extensible platform for tracking progress toward closing it.

Comments17 pages including appendices, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑