arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

套中之套:通过持续改进实现多日自主软件开发

Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement

Haoyang Yan, Min-le Su, Hangfan Zhang, Zhanhao Li, Chen Zhang, Shao Zhang, Yang Chen, Lei Bai, Shuyue Hu

arXiv 2609.01481首次发表:更新:

发表机构

Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Harness-of-Harness(HoH)框架,通过迭代规划-编码-测试循环实现自主软件开发的持续改进,在多基准测试中性能优于独立套具,还自主开发出具备完整功能的第一人称射击游戏。

AI 中文摘要

本文研究自主软件开发,其中基于大语言模型(LLM)的编码智能体可在无需人工干预的情况下,将高层需求转化为完整、功能齐全且可用的软件系统。我们提出Harness-of-Harness(HoH)框架,该框架使编码智能体能在自主开发过程中持续改进软件。HoH在现有编码智能体套具(harness)上运行,将其执行过程组织为迭代的规划-编码-测试循环。为在各循环间维持改进,HoH在修复与能力提升间取得平衡,将开发范围限定为小型可验证增量,将实现阶段的测试与独立评估分离,并约束可验证输出而非规定智能体工作流程。它逐步暴露交付物、角色特定工具及技能,鼓励复用而非重新创造,并维护版本化项目历史。在GameCraft-Bench、FrontierSWE和ProgramBench三个基准上,针对三套套具-模型组合(Codex搭配GPT-5.5、OpenCode搭配DeepSeek-V4-Pro、Pi搭配MiniMax-M3),HoH始终优于对应的独立套具,在三次迭代后平均相对增益达52.25%,最大增益达82.86%。在超过70次迭代的多日部署中,HoH自主开发出一款第一人称射击游戏,具备连贯剧情、完全实现的核心机制、可人工游玩的体验、精致的视觉效果及集成音频。

英文摘要

This paper studies autonomous software development, in which LLM-based coding agents transform high-level requirements into complete, functional, and usable software systems without human intervention. We introduce Harness-of-Harness (HoH), a framework that enables coding agents to continually improve software during autonomous development. HoH operates on existing coding-agent harnesses, and organizes their executions into iterative planning-coding-testing loops. To sustain improvement across loops, HoH balances repair with capability growth, scopes development into small and verifiable increments, separates implementation-time testing from independent evaluation, and constrains verifiable outputs rather than prescribing agent workflows. It progressively exposes deliverables, role-specific tools, and skills, encourages reuse rather than recreation, and maintains versioned project histories. On GameCraft-Bench, FrontierSWE, and ProgramBench, three harness-model pairs (Codex with GPT-5.5, OpenCode with DeepSeek-V4-Pro, and Pi with MiniMax-M3), HoH consistently outperforms the corresponding standalone harnesses, achieving an average relative gain of 52.25 percent and a maximum gain of 82.86 percent after three iterations. In a multi-day deployment with more than 70 iterations, HoH autonomously develops a first-person-shooter game, featuring a coherent storyline, fully implemented core mechanics, human-playable experience, polished visuals and integrated audio. Github: https://github.com/Flesymeb/HarnessOfHarness Project Page: https://flesymeb.github.io/HarnessOfHarness/

CommentsGithub: https://github.com/Flesymeb/HarnessOfHarness Project Page: https://flesymeb.github.io/HarnessOfHarness/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑