超越模型:揭秘软件工程智能体中的框架效应
Beyond the Model: Demystifying Harness Effects in Software Engineering Agents
浏览论文内容
中文总结 AI 辅助
本研究系统实证了软件工程智能体框架的影响,构建NanoHarness分析组件,发现框架效果依赖模型与任务,结构化工具和特定子代理提升稳定,组合框架可显著提升性能。
中文摘要 AI 辅助
基于大型语言模型(LLM)的智能体越来越多地用于软件工程任务,然而其性能并非仅由基础模型决定。智能体框架(harness)在很大程度上塑造了软件工程(SE)智能体与代码仓库交互、执行操作以及验证解决方案的方式。然而,框架设计的作用,尤其是在不同模型、任务和框架组件之间的差异,仍未得到充分理解。本文对SE智能体中的框架效应进行了系统的实证研究。我们首先使用来自两个著名开源权重模型系列(Qwen和DeepSeek)的十个模型,在三个基准测试(SWE-bench Pro、ProgramBench和GitTaskBench)上评估了两个代表性框架:mini-SWE-agent和OpenCode。随后,我们构建了NanoHarness,一个基于mini-SWE-agent的轻量级模块化框架,并利用它分析了五个代表性框架组件:工具注册表、上下文压缩、显式规划、子代理和惰性技能。实验结果表明,框架的有效性取决于模型能力和任务类型的共同作用。随着模型能力的提升,复杂框架在SWE风格的问题修复上提供的边际收益递减,但对于更强模型在更复杂和开放式的仓库级任务上可能有益。在ProgramBench上的组件级分析进一步表明,结构化工具使用和任务特定子代理提供了最稳定的改进,而上下文压缩和通用子代理可能损害仓库生成性能。组合使用时,NanoHarness在Qwen3.7-Max和DeepSeek-V4-Pro上分别比mini-SWE-agent提高了7.37和6.21个百分点,恢复了产品级框架的大部分增益。这些发现强调了框架设计是SE智能体性能的一等要素,并为构建更有效和高效的编码智能体提供了见解。
英文摘要
Large Language Model (LLM)-based agents are increasingly used for software engineering tasks, yet their performance is not determined by the base model alone. The agent harness substantially shapes how SE agents interact with repositories, execute actions, and validate solutions. However, the role of harness design remains insufficiently understood, especially across different models, tasks, and harness components. In this paper, we present a systematic empirical study of harness effects in SE agents. We first evaluate two representative harnesses, mini-SWE-agent and OpenCode, with ten models from two prominent open-weight model families, Qwen and DeepSeek, on three benchmarks: SWE-bench Pro, ProgramBench, and GitTaskBench. We then construct NanoHarness, a lightweight modular harness built on top of mini-SWE-agent, and use it to analyze five representative harness components: tool registry, context compression, explicit planning, subagents, and lazy skills. Experimental results show that harness effectiveness depends jointly on model capability and task type. Complex harnesses provide diminishing marginal gains on SWE-style issue repair as model capability improves, but can benefit stronger models on more complex and open-ended repository-level tasks. Component-level analysis on ProgramBench further shows that structured tool use and task-specific subagents provide the most stable improvements, while context compression and general subagents can hurt repository-generation performance. When combined, NanoHarness improves over mini-SWE-agent by 7.37 and 6.21 percentage points on Qwen3.7-Max and DeepSeek-V4-Pro, respectively, recovering most of the gains of product-level harnesses. These findings highlight harness design as a first-class factor in SE-agent performance and provide insights for building more effective and efficient coding agents.
发表机构
- Nanjing University of Science and Technology(南京理工大学)
- Technical University of Munich(慕尼黑工业大学)
- Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。