arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19134cs.CLcs.CY

ScienceIDE:将全球科学代码库转化为智能体可学习环境

ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments

Hejia Geng, Zesen Huang, Haoyang Li, Wenbin Li, Koutian Wu, Zihan Zhou, Yuanbo Pang, Weihao Liu, Zigong Xu, Zhiping Li, Zongzheng Zhang, Chuanfei Dong, Jiankai … 展开作者

Hejia Geng, Zesen Huang, Haoyang Li, Wenbin Li, Koutian Wu, Zihan Zhou, Yuanbo Pang, Weihao Liu, Zigong Xu, Zhiping Li, Zongzheng Zhang, Chuanfei Dong, Jiankai Sun, Tianzhe Zheng, Fengyu Xie, Yue Ma, Yueheng Shi, Tong Xie, Zonglin Di, Xianrong Liu, Qucheng Gao, Yimin Liu, Jiaming Pan, Sheng Huang, Xiao-Han Ma, Lanqing Yuan, Zhenlin Zhu, Ziang Liu, Ziyang Xu, Junkai Wang, Kangkai Liang, Jiayi Xian, Zehong Zhao, Liuwei Xu, Jingxu Xie, Peijin Zhang, Qiang Gao, Chengyi Xing, Zhe Zhao, Xi Wang, Yaopeng Xing, Xing Meng, Zhenfei Yin, Yingcheng Wu, Ling Yang

首次发表
浏览论文内容

中文总结 AI 辅助

针对科学代码难以转化为学习体验的瓶颈,提出ScienceIDE基础设施,将科学代码仓库转化为可执行环境,训练PhAI-IDE系列模型,在代码修复和通用基准上取得提升。

中文摘要 AI 辅助

科学代码仓库以可执行的模型、方法和工具形式编码了数十年来的人类知识。然而,碎片化的工具链、隐含的领域约定和专门的正确性标准使得这些知识难以转化为可靠的学习体验——我们将这一挑战称为科学经验瓶颈。我们提出ScienceIDE,一种将全球科学代码转化为科学智能体可编程环境的基础设施。在专家定义的科学案例和验收标准的指导下,智能体将代码仓库转化为支持任务生成、执行和科学验证的可执行环境。这些环境为监督微调、强化学习和评估提供了共享基础。利用经过验证的交互轨迹,我们训练了PhAI-IDE-72B、PhAI-IDE-9B和PhAI-IDE-4B。该模型系列在保留的科学代码修复任务以及选定的通用基准(涵盖代码、推理和知识)上表现出提升,为科学经验向更广泛能力的正迁移提供了证据。ScienceIDE为智能体学习和科学实践的集成工作空间奠定了基础,使人类的科学软件成为发展科学智能的共享基础。代码:此https URL

英文摘要

Scientific code repositories encode decades of human knowledge in executable models, methods, and tools. Yet fragmented toolchains, implicit domain conventions, and specialized correctness criteria make this knowledge difficult to convert into reliable learning experience-a challenge we call the scientific experience bottleneck. We introduce ScienceIDE, infrastructure for turning the world's scientific code into programmable environments for scientific agents. Guided by expert-defined scientific cases and acceptance criteria, agents transform repositories into executable environments that support task generation, execution, and scientific verification. These environments provide a shared foundation for supervised fine-tuning, reinforcement learning, and evaluation. Using verified interaction trajectories, we train PhAI-IDE-72B, PhAI-IDE-9B, and PhAI-IDE-4B. The model family shows gains in held-out scientific-code repair and across selected general-purpose benchmarks in code, reasoning, and knowledge, providing evidence of positive transfer from scientific experience to broader capabilities. ScienceIDE lays the foundation for an integrated workspace for agent learning and scientific practice, making humanity's scientific software a shared substrate for developing scientific intelligence. Code: https://github.com/aitofound/ScienceIDE

发表机构

  • PhAI-Labs(PhAI实验室)
  • Qwen

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑