OpenAl4S:代码即行动,科学即会话
OpenAI4S: Code as Action, Science as Sessions
- Peking University Shenzhen Graduate School(北京大学深圳研究生院)
- Tsinghua-Peking Joint Center for Life Sciences(清华-北大生命科学联合中心)
- Tsinghua University(清华大学)
- Beijing Yuankong Intelligent Technology Co., Ltd.(北京元空智能科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
OpenAI4S 基于“代码即行动,科学即会话”原则,通过持久化运行时和会话管理提升 AI 科学研究的可靠性,在 36 个场景中得分 7.83,优于通用编码框架。
AI中文摘要:
AI 合作科学家(AI co-scientists)有望加速计算研究,但在长期研究中,工作流程还必须保持可检查性、可恢复性和可重现性,这需要持久化的计算状态和来源追踪。在此,我们提出 OpenAI4S,一个基于“代码即行动,科学即会话”原则构建的开源科学研究智能体。OpenAI4S 结合了持久化计算运行时与研究会话管理:编排通过结构化工具调用处理,而科学行动则表示为在持久化 Python 和 R 内核中执行的完整代码单元。一个仅追加的操作账本(Action Ledger)、逐单元执行记录、版本化工件、环境记录和工作区检查点共同保留了结果的产生过程,并支持会话恢复、分支和扩展。可配置的沙箱、权限控制以及代码和轨迹筛选提供了补充性安全保障。我们在涵盖逆合成、分子动力学、蛋白结合剂设计、蛋白质突变、催化剂筛选和矿物光谱学的 36 个研究场景上评估了 OpenAI4S,测量了科学任务准确性、工作流程完整性和所生成仓库的可重现性。OpenAI4S 取得了 7.83 的总体得分,而使用三个前沿模型评估的通用编码框架得分为 5.7–6.4,其中在长周期和计算密集型工作流程上提升最大。这些结果表明,将持久化执行与会话级来源追踪相结合可以提高 AI 辅助科学工作流程的可靠性。对于所有被评估的系统(包括我们自己的),环境规范化和完全可重运行性仍然薄弱,因此可重现性对于科学智能体而言仍是一个开放问题。该系统以 MIT 许可证在 https://this https URL 上提供。
英文摘要:
AI co-scientists could accelerate computational research, but over a long-running study the workflow also has to stay inspectable, resumable and reproducible, which requires persistent computational state and provenance. Here we present OpenAI4S, an open-source scientific research agent built around the principle of \emph{Code as Action, Science as Sessions}. OpenAI4S combines a persistent computing runtime with research-session management: orchestration is handled through structured tool calls, while scientific actions are represented as complete code cells executed in persistent Python and R kernels. An append-only Action Ledger, per-cell execution records, versioned artifacts, environment records, and workspace checkpoints preserve how results were produced and support session recovery, branching, and extension. Configurable sandboxing, permission controls, and code and trajectory screening provide complementary safeguards. We evaluate OpenAI4S on 36 research scenarios spanning retrosynthesis, molecular dynamics, protein binder design, protein mutation, catalyst screening, and mineral spectroscopy, measuring scientific task accuracy, workflow completeness, and reproducibility of the resulting repositories. OpenAI4S achieves an overall score of 7.83, compared with 5.7--6.4 for a general-purpose coding harness evaluated with three frontier models, with the largest gains on long-horizon and computation-intensive workflows. These results suggest that integrating persistent execution with session-level provenance can improve the reliability of AI-assisted scientific workflows. Environment specification and full rerunnability remain weak for every evaluated system, ours included, so reproducibility is still an open problem for scientific agents. The system is available under the MIT license at \href{https://github.com/PKU-YuanGroup/OpenAI4S}{github.com/PKU-YuanGroup/OpenAI4S}.