arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34378cs.CV

Marathoner:超长时域自主智能

Marathoner: Ultra-Long-Horizon Autonomous Intelligence

Zhang Ruiyang, Ou Jinpeng, Xie Yifan, Zhou Jingang, Pan Lirui, Guo Qingpei, Zheng Zhedong

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出Marathoner,一种通过后训练流程(任务合成、拒绝采样微调与强化学习)实现超长时域自主执行的智能体模型,在5个基准上超越基础模型及强专有模型,可连续工作10小时以上。

中文摘要 AI 辅助

人类天生具备为长期目标持续工作的能力。面对一项具有挑战性的任务,人类可以连续工作数月甚至数年以实现特定目标。在本文中,我们提出了Marathoner,一种具备超长时域执行能力的自主智能体模型。具体而言,我们提出了一套全面的后训练流程,以将这一关键能力注入基础模型。对于超长时域任务合成,我们利用来自多样化GitHub仓库的包含1000行以上新代码的主要发布拉取请求,作为合成具有挑战性的任务级数据的主要来源。此外,我们引入了多任务链式(Multi-Task Chaining),将多个生成的任务链接为单个更具挑战性的任务,从而能够合成具有前沿难度的任务。对于拒绝采样微调,我们将强教师模型与多样化测试工具相结合,以在合成任务上生成轨迹,并对基础模型使用拒绝采样的轨迹进行监督微调。对于强化学习,冷启动模型在回放过程中通过独立沙箱中的测试工具执行真实世界操作,有效促进获取真正的超长时域执行能力。我们进一步提出了一种新颖的奖励策略,即后期阶段奖励(Later Stage Bonus Reward),该策略明确鼓励模型在执行后期阶段进行有意义的操作。通过在包含超长时域任务的5个基准上进行广泛评估,Marathoner相对于基础模型取得了持续且显著的性能提升,甚至超越了强专有模型的性能。进一步分析表明,Marathoner能够在极具挑战性的任务上持续工作10小时以上,并进行1000次以上的工具调用。

英文摘要

Humans naturally possess the ability to work persistently toward long-term goals. Given a challenging task, humans can continuously work for months or even years to accomplish a specific objective. In this paper, we propose Marathoner, an autonomous agentic model possessing the ability of ultra-long-horizon execution. Specifically, we propose a comprehensive post-training pipeline to instill this critical capability into base model. For Ultra-Long-Horizon Task Synthesis, we leverage major release PRs containing 1000+ lines of new code from diverse GitHub repositories as the primary source for synthesizing challenging task-level data. Additionally, we introduce Multi-Task Chaining, which chains multiple generated tasks into a single more challenging task, enabling the synthesis of tasks with frontier-level difficulty. For rejection sampling finetuning, we combine strong teacher model with diverse harnesses to generate trajectories on our synthesized tasks and conduct supervised finetuning on base model with rejection sampled trajectories. For reinforcement learning, cold-started model performs real-world execution through harnesses in independent sandboxes during rollout process, effectively facilitating the acquisition of genuine ultra-long-horizon execution capability. We further propose a novel reward strategy, Later Stage Bonus Reward, which explicitly encourages model to perform meaningful maneuvers during later stages of execution. Through extensive evaluation on 5 benchmarks containing ultra-long-horizon tasks, Marathoner achieves consistent and substantial performance improvements over base model and even surpasses performance of strong proprietary model. Further analysis shows that Marathoner can consistently work for 10+ hours and conduct 1000+ tool calls on highly challenging tasks.

发表机构

  • University of Macau(澳门大学)
  • Ant Group(蚂蚁集团)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑