arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15579cs.SEcs.AIcs.ETcs.PL

Kozuchi智能体:一款与语言无关的开源权重软件修复智能体

Kozuchi Agent: A Language-Agnostic Open-Weight Agent for Software Repair

Mehdi Bahrami, Kosaku Kimura, Satoshi Munakata, Satoshi Nakashima, Yu Ishikawa, Kosuke Maeda, Nao Soma, Kenichi Kobayashi, Keisuke Miyazaki, Keizo Kato, Shigeki… 展开作者

Mehdi Bahrami, Kosaku Kimura, Satoshi Munakata, Satoshi Nakashima, Yu Ishikawa, Kosuke Maeda, Nao Soma, Kenichi Kobayashi, Keisuke Miyazaki, Keizo Kato, Shigeki Fukuta, Tatsuo Kumano, Nobutaka Imamura, Kevin Musgrave, Shahbaz Abdul Khader, Kwun Ho Ngan, Joe Townsend, Fayas Asharindavida, Matthieu Parizy, Akira Sakai, Yuma Ichikawa, Yang Zhao, Michiaki Takizawa, Taku Fukui, Hiroki Ohtsuji, Wei-Peng Chen, Hiromichi Kobashi

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出Kozuchi智能体,一款与语言无关的开源权重软件修复智能体,结合CI评估流水线,在SWE-bench等基准上取得优异性能,降低操作接触点,为软件修复提供新方案。

中文摘要 AI 辅助

工业软件工程团队日益需要大型语言模型(LLM)智能体,将缺陷报告转化为正确的补丁,但基准规模的操作会带来长 horizon、工具使用规范性、上下文持久性、异构集群和评估复用等问题。我们提出了Kozuchi智能体,这是一款与语言无关的开源权重修复智能体,以及由持续集成(CI)运行的评估流水线。其包含显式阶段、持久状态、确定性工具、与模型无关的动作接口和跨智能体测试时选择,使运行过程可审计且可重复。使用本地部署的Qwen3.5-27B模型,无需微调,采用TTS@8设置,Kozuchi在官方评估器上解决了500个SWE-bench Verified实例中的374个。在Multi-SWE-bench Java数据集上性能未变,这款相同的270亿参数智能体解决了128个实例中的41个(32.03%),在严格开源权重提交中排名第一,在全部42个提交中排名第四;在Python数据集上,在135个提交中排名第12,在开源权重系统中排名第一。各阶段行为在不同语言间的差异保持在±5个百分点以内。剩余失败主要反映语义正确性、Java特定的测试 harness 问题和选择错误。在两个赛道中,按参数数量与开源/本地同行相比,结果表现更优。对候选多样性、选择器遗憾和补丁可靠性的分析显示,剩余差距主要在于语义正确性和选择,而非编辑格式或专有模型访问。在操作层面,可复用的CI阶段将异构内部集群的操作员接触点从5个减少到1个。

英文摘要

Industrial software-engineering teams increasingly need LLM agents that turn bug reports into correct patches, yet benchmark-scale operation adds long horizons, tool-use discipline, context persistence, heterogeneous clusters, and evaluation reuse. We present Kozuchi Agent, a language-agnostic open-weight repair agent and CI-operated evaluation pipeline. Explicit phases, persistent state, deterministic tools, a model-independent action interface, and cross-agent test-time selection make runs auditable and repeatable. With locally hosted Qwen3.5-27B, no fine-tuning, and TTS@8, Kozuchi resolves 374/500 SWE-bench Verified instances on the official evaluator. Unchanged on Multi-SWE-bench Java, the same 27-billion-parameter agent resolves 41/128 instances (32.03%), ranking first among strict open-weight submissions and fourth of 42 overall; on Python it ranks 12th of 135 and first among open-weight systems. Per-phase behavior remains within +/-5 percentage points across languages. Remaining failures mainly reflect semantic correctness, Java-specific harness issues, and selection errors. Across both tracks, results compare favorably with open/local peers by parameter count. Analysis of candidate diversity, selector regret, and patch reliability shows that the remaining gap is primarily semantic correctness and selection rather than edit formatting or proprietary-model access. Operationally, reusable CI stages reduce operator touch-points from five to one across heterogeneous internal clusters.

发表机构

  • Fujitsu Research of America(美国富士通研究所)
  • Fujitsu Research(富士通研究所)
  • Fujitsu Research of Europe(欧洲富士通研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑