发表机构
NVIDIA; UCLA; MIT; Tsinghua University(英伟达; 加州大学洛杉矶分校; 麻省理工学院; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Humanize,一种基于判断工程的智能体编码多智能体工作流,通过机械门控和独立审查提升完成可靠性,在多项基准和竞赛中取得领先成绩。
AI 中文摘要
智能体编码使得代码生成变得廉价,但可靠的完成仍然困难:编写代码的智能体是判断其是否完成的薄弱评审者。我们提出Humanize,一个围绕判断工程构建的智能体编码多智能体编排工作流:在规划、实现、审查和学习之间的边界上,进行显式的、机械强制执行的决策。人类批准一份计划合同,一个构建者智能体分轮实现它,一个来自另一供应商的审查者智能体决定完成;确定性钩子(而非模型)在这些角色之间路由工作并强制执行72个机械门控。将其视为仓库状态上的马尔可夫链,交替的构建者和审查者从两个模型联合采样,因此缺陷只有在两者都遗漏时才能存活。我们通过其部署、118个真实循环的公开事后分析及其应用来研究Humanize。在108天的68个版本中,它获得了1,468个GitHub星标。应用包括在上游审查下进行的567文件gem5构建系统迁移;内核设计智能体,它通过内核知识库和分析反馈扩展循环,并在MLSys 2026 FlashInfer竞赛的所有三个全智能体赛道中进入前三名;以及通过Humanize奥林匹克智能体(HOA),在IOI 2026、IMO 2026、IPhO 2026和IBO 2024中获得满分,在IChO 2026中获得418.5/437(金牌)。Humanize还在PutnamBench上达到672/672,并在Lean-Eval排行榜上排名第一(251/303),甚至与专业数学家竞争。事后分析表明,独立审查能捕捉到无根据的构建者声明,但停止仍然是一个关键弱点。在按阶段分轮的报告里,三分之二的轮次发生在实现被接受之后。这一证据是观察性的,而非工作流的受控比较。
英文摘要
Agentic coding makes code generation cheap, but reliable completion remains difficult: the agent that writes the code is a weak judge of whether it is done. We present Humanize, a multi-agent orchestration workflow for agentic coding built around judgement engineering: explicit, mechanically enforced decisions at the boundaries between planning, implementation, review, and learning. A human approves a plan contract, a builder agent implements it in rounds, and a reviewer agent from another vendor decides completion; deterministic hooks, not a model, route work between these roles and enforce 72 mechanical gates. Viewed as a Markov chain over repository states, alternating builder and reviewer samples jointly from two models, so a defect survives only if both miss it. We study Humanize through its deployment, 118 public postmortems of real loops, and its applications. Over 68 versions in 108 days, it gathered 1,468 GitHub stars. Applications include a 567-file gem5 build-system migration under upstream review; Kernel Design Agents, which extend the loop with a kernel knowledge base and profiling feedback and placed in the top three of all three Full-Agent tracks of the MLSys 2026 FlashInfer contest; and, through Humanize Olympiad Agents (HOA), full scores in IOI 2026, IMO 2026, IPhO 2026, and IBO 2024, 418.5/437 in IChO 2026 (gold-medal). Humanize also achieves 672/672 on PutnamBench and ranks first (251/303) on Lean-Eval's leaderboard even competiting with professional mathematicians. The postmortems show that independent review catches unsupported builder claims, but stopping remains a key weakness. In reports that separate rounds by phase, two thirds of rounds occurred after implementation was accepted. This evidence is observational, not a controlled comparison of workflows.