arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

openJiuwen:超越静态框架的长程编码智能体

openJiuwen: Beyond Static Harnesses for Long-Horizon Coding Agents

openJiuwen Team, Tao Yu, Xinyu Zhang, Qianqian Chen, Xiaoneng Xiang, Chia Kwangyang, Xingchen Huang, Ran Chen, Yangkai Ding, Zheng Wang, Yeo Boon Hong, Bingzheng Gan, Enrui Hu, Shuo Cheng, Deyang Li, Ruifeng Shi, Hongbo Wang, Qi Ye, Xuefeng Jin, Zhangchun Zhao

arXiv 2608.27969首次发表:更新:

发表机构

Huawei Technologies Co., Ltd.(华为技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

openJiuwen是一款开源智能体框架,针对长程编码智能体的结构可组合性与运行时适应性挑战设计,在SWE-bench Verified和Terminal-Bench 2.1上的准确率分别达82.6%、87.19%,性能优于官方排行榜顶尖方案。

AI 中文摘要

长程编码智能体在不断演化的仓库状态上运行,且越来越依赖异构能力、委派智能体和多智能体协作。这些趋势给智能体框架带来两个互补挑战:其一,开发者需组合能力、重配置执行逻辑、扩展日益复杂的智能体系统,无需反复重构编排;其二,复杂编码任务会持续产生新证据,如语义诊断、执行结果、任务进度及不断变化的上下文相关性,这些应动态影响后续运行时决策。我们将这些挑战定义为结构可组合性与运行时适应性。我们提出openJiuwen,一个专为开发者可组合性和自适应任务执行设计的开源框架。openJiuwen提供共享执行基础,以及基于Rail的能力组合,涵盖单一智能体、委派子智能体和Swarm Flow,使开发者能在通用执行语义下构建复杂智能体框架。它还围绕固定模型策略调整框架控制的运行时决策,让不断演化的证据动态影响上下文、反馈和任务控制,以确保成功完成任务。我们在SWE-bench Verified和Terminal-Bench 2.1上对openJiuwen进行系统评估,其分别达到82.6%和87.19%的准确率,超出官方排行榜最强选中点估计值3.4和3.39个百分点。这些结果表明,openJiuwen在复杂编码任务上表现出色,同时提供可组合且自适应的框架设计。

英文摘要

Long-horizon coding agents operate over evolving repository states while increasingly relying on heterogeneous capabilities, delegated agents, and multi-agent coordination. These trends pose two complementary challenges for the agent harness. First, developers need to compose capabilities, reconfigure execution logic, and scale increasingly complex agent systems without repeatedly rebuilding orchestration. Second, complex coding tasks continuously produce new evidence---such as semantic diagnostics, execution outcomes, task progress, and changing context relevance---that should dynamically influence subsequent runtime decisions. We characterize these challenges as Structural Composability and Runtime Adaptivity. We present openJiuwen, an open-source harness designed for both developer composability and adaptive task execution. openJiuwen provides a shared execution substrate and Rail-based capability composition across single agents, delegated sub-agents, and Swarm Flow, enabling developers to construct sophisticated agent harnesses under common execution semantics. It further adapts framework-controlled runtime decisions around a fixed model policy, allowing evolving evidence to dynamically affect context, feedback, and task control toward successful completion. We systematically evaluate openJiuwen on SWE-bench Verified and Terminal-Bench 2.1, where it achieves 82.6% and 87.19%, respectively, exceeding the strongest selected official-leaderboard point estimates by 3.4 and 3.39 percentage points. These results show that openJiuwen achieves strong performance on complex coding tasks while providing a composable and adaptive harness design.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑