arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23552cs.AIcs.CLcs.SE

Prime Agent:一种自改进的RLM管控框架

Prime Agent: A Self-Improving RLM Harness

  • Princeton University(普林斯顿大学)
  • MIT(麻省理工学院)
  • Prime Intellect

机构由 AI 辅助整理,请以论文原文为准。

Seth Karten, Alex L. Zhang, Kevin Thomas, Sebastian Müller, Elie Bakouch, Daniel Auras, Mika Senghaas, Fares Obeid, Konstantin Dunas, Johannes Hagemann, Sami Jaghouar

AI总结:

Prime Agent是一款开源RLM管控框架,通过标准化流程提升模型长周期能力,在ARC-AGI-3等任务中大幅提升性能,支持编码等工作流与并行化任务。

AI中文摘要:

语言模型是顺序处理器,但长周期智能体需要超出模型权重和活跃上下文的外部信息与计算资源。Prime Agent是一款用于长周期评估和编码智能体工作流的开源管控框架。持续运行的IPython REPL遵循递归语言模型(Recursive Language Model,RLM)抽象,用于程序化上下文处理和测试时计算;而持续管控框架(Continual Harness)则在不同轨迹间保留历史记录、记忆、技能、提示词和子智能体规格。递归子智能体通过直接的智能体间通信进行协调,智能体视图(Agents View)允许人类检查和管理由守护进程支持的会话。Prime Agent标准化了执行、恢复、验证和资源核算流程,同时将策略构建工作交给模型完成。这种低摩擦、高表达性的机制可防止管控框架故障演变为模型故障,并将评测导向模型真实的最大潜在能力。Prime Agent将ARC-AGI-3 RHAE的Best@1指标从30%提升至95.5%,在长上下文编码、GPU内核生成、模拟器构建和自主nanoGPT速跑任务中,其性能与原生及主流管控框架相当或更优。在Factorio游戏任务中,研究发现细化机制可实现持续技术进步,专用子智能体则支持并行化工作。代码可在指定URL获取。

英文摘要:

Language models are sequential processors, but long-horizon agency requires external information and computation beyond model weights and active context. Prime Agent is an open-source harness for long-horizon evaluation and coding-agent workflows. A persistent IPython REPL follows the Recursive Language Model abstraction for programmatic context processing and test-time compute, while Continual Harness preserves histories, memories, skills, prompts, and subagent specifications across trajectories. Recursive subagents coordinate through direct agent-to-agent communication, and the Agents View lets humans inspect and manage daemon-backed sessions. Prime Agent standardizes execution, recovery, verification, and resource accounting while leaving strategy construction to the model. This low-friction, expressive membrane prevents harness failures from becoming model failures and pushes measurement toward the model's true maximal underlying capability. Prime Agent raises ARC-AGI-3 RHAE Best@1 from 30% to 95.5% and matches or exceeds native and popular harnesses across long-context coding, GPU-kernel generation, emulator construction, and autonomous nanoGPT speedruns. On Factorio, we find refinement allows for continuous technology progression and dedicated subagents enable parallelized work. Code is available at https://github.com/PrimeIntellect-ai/prime-agent.

补充信息

↑