arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14354cs.AI

ScienceFlow:面向机器学习研究、科学发现及更广泛领域的长时智能体

ScienceFlow: A long-horizon agent for ML research, scientific discovery and beyond

Mingming Zhao, Jiqian Dong, Kangping Xu, Zadid Hasan, Chengrui Fan, Shan Jiang, Shuai Mao, Yating Ling, Linyi Zou, Tailin Zhou, Yun Hin Chan, Wenkai Zhang, Zhan… 展开作者

Mingming Zhao, Jiqian Dong, Kangping Xu, Zadid Hasan, Chengrui Fan, Shan Jiang, Shuai Mao, Yating Ling, Linyi Zou, Tailin Zhou, Yun Hin Chan, Wenkai Zhang, Zhanhong Zhou, Guowei Huang, Hongliang Li, Wenjing Cun, Zhitang Chen, Mingxuan Yuan, Yanhui Geng

首次发表
浏览论文内容

中文总结 AI 辅助

ScienceFlow是一款端到端自主研究智能体框架,通过ESTRA机制与证据感知执行控制器实现长时研究的状态管理与资源分配,在24小时预算内的MLE-bench基准测试中取得70.22%的Any-Medal最优分数,为长时自主研究提供了有效方案。

中文摘要 AI 辅助

让大语言模型(LLM)智能体在长时范围内维持高效、稳定且与目标对齐的研究过程,是自主机器学习与科学发现的核心挑战,因为研究进展依赖于持续管理不断演变的状态、探索决策及计算资源。尽管开创性的自主研究智能体已取得显著成果,但仍缺乏连续性机制、死胡同恢复机制以及基于价值的计算资源分配机制,这会固有地损害整体搜索效率、浪费计算资源并降低最终成功的概率。为填补这一空白,我们提出ScienceFlow,这是一个端到端的自主研究智能体框架,它将长时研究工作组织为基于可执行工作空间的研究片段,将研究进展表示为可恢复的可执行状态,从而支持高效的探索、修订与执行。研究片段之间的转换由通过重新锚定实现的可执行状态转换(ESTRA)控制,该机制会选择当前活跃状态或存档状态作为下一个锚点,并确定是继续还是重定向研究轨迹。一个感知证据的执行控制器会根据资源可用性、剩余预算及已验证的进展,将资源分配给实际作业。我们在涵盖机器学习、科学建模和数学优化的任务上对ScienceFlow进行了评估。在各类长时基准测试上的结果表明,ScienceFlow能够维持有效的研究过程,其中最突出的是在24小时预算内,其在完整的MLE-bench上取得了70.22%的Any-Medal分数,达到了当前最优(SOTA)水平,比之前报告的结果高出4.92个百分点。ScienceFlow的功效进一步证明,高效的状态管理、自适应探索以及与目标对齐的执行,对于将自主研究扩展到超出短时交互的范围至关重要。

英文摘要

Enabling LLM agents to sustain productive, stable, and goal-aligned research over extended horizons is a central challenge for autonomous machine learning and scientific discovery, as progress hinges on continuously managing evolving state, exploration decisions, and computational resources. Pioneering autoresearch agents, despite great success, still lack mechanisms for continuity, recovery from dead ends, and value-driven compute allocation, which inherently undermines overall search efficiency, wastes computational resources, and lowers the chance of ultimate success. To bridge this gap, we introduce ScienceFlow, an end-to-end autoresearch agent framework that organizes long-horizon research work into research segments grounded in executable workspaces. It represents research progress as recoverable executable states, enabling efficient exploration, revision, and execution. Transitions between research segments are governed by Executable-State Transition through Re-Anchoring (ESTRA), which selects either the live state or an archived state as the next anchor and determines whether to continue or redirect the research trajectory. An evidence-aware execution controller allocates resources to physical jobs based on resource availability, remaining budget, and validated progress. We evaluate ScienceFlow on tasks spanning machine learning, scientific modeling, and mathematical optimization. Results on diverse long-horizon benchmarks demonstrate its ability to sustain effective research processes, highlighted by a SOTA 70.22 percent Any-Medal score on the full MLE-bench within a 24-hour budget, outperforming prior reported results by 4.92 percentage points. The efficacy of ScienceFlow further demonstrates that efficient state management, adaptive exploration, and objective-aligned execution are critical for scaling autonomous research beyond short-horizon interactions.

发表机构

  • Noah’s Ark Lab, Huawei(华为诺亚方舟实验室)

机构由 AI 辅助整理,请以论文原文为准。

相关深度报道

↑