发表机构
Microsoft; Shanghai Jiao Tong University; Fudan University; Nanjing University; Tsinghua University; The University of Hong Kong; Peking University; The Chinese University of Hong Kong, Shenzhen; Donghua University(微软公司; 上海交通大学; 复旦大学; 南京大学; 清华大学; 香港大学; 北京大学; 香港中文大学(深圳); 东华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Argus是面向长程推理的通用智能体运行时,通过多角色协作与自进化,在多个基准测试中表现优于同类方法,还完成了实际应用任务并产出结构化轨迹。
AI 中文摘要
长程推理需要一种智能体运行时,当证据支持当前方法时能持续运行,而当测量结果显示失败、隐藏约束或目标设定错误时能调整方向。我们提出了Argus,这是一种持久、自进化的运行时,其中Manager(管理者)、Planner(规划者)、Engineer(工程师)和Reviewer(审核者)在持久的项目状态上执行有界任务。Argus将稳定的用户意图与操作目标、约束及验证标准分离,仅在角色所有的审核且(若可用)任务原生验证后,才允许记忆、技能、程序、验证器、路由决策及被拒绝的路由存在。模型权重保持固定;自进化通过持久的运行时状态和控制策略实现,在操作者所有的升级点之间自主执行。在7个GPT-5.5基准领域中,Argus在SWE-Bench Pro上达到约78%,而Direct Copilot为59%,同时使用1.41倍的总token量。在经过验证门控的自进化后,成熟的SWE-Bench批次相比初始批次,每个任务的求解输入token量减少21%,主动工作流程时间减少15%,同时记录了34次验证器恢复和22次严格审核循环的挽救。Argus在AARRI-Bench上达到76.8%,在数学数据合成上有28.0分的差距,并有具有竞争力的GPU内核和语言模型训练结果。除基准测试外,一个优化的RWKV6内核已被合并到上游;一个为期多天的数学活动保留了被伪造的路由和有证明支持的前沿更新;六个论文流水线完成了254个任务,其中有16次阶段回滚。这些结果表明,固定权重、自进化的管控工具可以修改、恢复并积累已验证的方法,同时为未来的监督学习和强化学习生成结构化轨迹。
英文摘要
Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements reveal failure, hidden constraints, or a misspecified objective. We present Argus, a persistent, self-evolving runtime in which Manager, Planner, Engineer, and Reviewer execute bounded missions over durable project state. Argus separates stable user intent from operational objectives, constraints, and verification criteria, and admits memories, skills, procedures, verifiers, routing decisions, and rejected routes only after role-owned review and, when available, task-native verification. Model weights remain fixed; self-evolution occurs through persistent runtime state and control policy, with autonomous execution between operator-owned escalation points. Across seven GPT-5.5 benchmark arenas, Argus achieves about 78% on SWE-Bench Pro versus 59% for Direct Copilot while using 1.41 times the aggregate tokens. After verification-gated self-evolution, mature SWE-Bench waves use 21% fewer solve-input tokens and 15% less active workflow time per task than startup waves, while recording 34 verifier recoveries and 22 strict review-loop rescues. Argus also reaches 76.8% on AARRI-Bench and a 28.0-point gap on mathematical data synthesis, with competitive GPU-kernel and language-model-training results. Beyond benchmarks, an optimized RWKV6 kernel was merged upstream; a multi-day mathematics campaign retained falsified routes and proof-backed frontier updates; and six paper pipelines completed 254 missions with 16 stage rollbacks. These results show that a fixed-weight, self-evolving harness can revise, recover, and accumulate verified approaches while producing structured trajectories for future supervised and reinforcement learning.