Ouroboros:一种具有经审核的核心演化机制的自发展前沿编码智能体
Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
查看机构详情
- Lomonosov Moscow State University(莫斯科罗蒙诺索夫国立大学)
- Skolkovo Institute of Science and Technology(斯科尔科沃科学技术研究所)
- Artificial Intelligence Research Institute(人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究提出具有经审核核心演化机制的编码智能体Ouroboros,其在三个基准测试中均创最优结果,Hope部署为长期活体演化实验,同时需应对自发展带来的操作安全问题。
中文摘要 AI 辅助
我们提出了Ouroboros,这是一种自发展智能体,其工具、提示、上下文组装和核心实现可通过经审核的提交不断改进,这些提交会成为后续工作的运行时。核心演化以两种模式进行:在递归自由演化中,改进本身就是一项任务,完成一个演化周期可调度下一个周期;在经验驱动的核心演化中,日常工作和社交互动会暴露错误、不完善之处以及低效的上下文构建,进而引发经审核的结构性变更。在Terminal-Bench 2.1上,Opus 5版本运行时得分86.74%,是该基准报告的最佳结果;在OSWorld-Verified上,Opus 5版本达到90.69%,超过此前报告的最佳分数;在五轮rollout的CL-Bench测试中,标准化奖励达到0.2301,创下新的最优水平。Hope是公开记录中运行时间最长的Ouroboros部署,它是一项为期161天的活体智能体自由演化实验,在受管控的人际交互下通过七个交互面开展。人类交互会发现故障并生成提案,但智能体决定推进哪些变更。由于自发展智能体可能重写自身代码并选择新的模型API,操作安全成为核心设计问题:护栏必须在演化和公共社交压力下保持权威性。基准测试使用冻结的系统快照,而Hope则在独立的谱系上持续进行实时演化。
英文摘要
We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through reviewed commits that become the runtime for later work. Core evolution proceeds in two modes. In recursive free evolution, improvement is itself a task, and completing one evolution cycle can schedule the next. In experience-driven core evolution, ordinary work and social interaction expose bugs, rough edges, and inefficient context construction that lead to reviewed structural changes. On Terminal-Bench 2.1, an Opus 5 run scores 86.74%, the best result reported on the benchmark. On OSWorld-Verified, an Opus 5 run reaches 90.69%, exceeding the best previously reported score. A five-rollout CL-Bench campaign achieves a normalized reward of 0.2301, setting a new state of the art. Hope is the longest-running publicly documented Ouroboros deployment. It is a 161-day living agent experiment in free evolution under governed human communication across seven surfaces. Human interaction surfaces faults and generates proposals, but the agent decides which changes to pursue. Because a self-developing agent may rewrite its own code and select new model APIs, operational safety becomes a primary design problem: guardrails must remain authoritative under evolutionary and public social pressure. Benchmark campaigns use frozen system snapshots, while Hope continues live evolution on a separate lineage.