arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07832cs.SEcs.AIcs.CLcs.MA

面向软件工程的模块化可执行开发原语(Dev-Primitives)的驾驭工程(Harness Engineering)

Harness Engineering for Software Engineering via Modular Executable Dev-Primitives

Haibo Jin, Xinjie Li, Peng Kuang, Haohan Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM智能体在长周期软件工程任务中的上下文爆炸与语义漂移问题,提出模块化可执行Dev-Primitives抽象及HERMES框架,通过依赖感知激活与缺陷诊断机制,在四个基准上平均提升12.4%,并显著降低推理成本。

中文摘要 AI 辅助

配备终端访问能力的大型语言模型(LLM)在自动化软件工程任务中展现出强大能力。然而,现有智能体在长周期工作流中仍然脆弱,它们必须反复重建分散在源文件、配置、测试、依赖项和运行时行为中的程序状态,导致交互历史越来越长、上下文爆炸和语义漂移。大型代码库进一步增加了识别任务相关组件的难度。为解决这些挑战,我们引入了Dev-Primitives(开发原语),一种模块化且可执行的抽象,将仓库组件从被动的软件工件转变为软件工程中的主动参与者。每个Dev-Primitive将一个仓库工件与一个驻留LLM配对,使该工件拥有一个基于其自身实现和依赖关系的智能体原生接口,从而实现自然语言推理、组件间通信和局部自我修改。基于Dev-Primitives,我们提出了HERMES,一个通过模块化可执行开发原语实现软件工程驾驭(Harness Engineering)的框架,该框架通过依赖感知的动态激活机制和将执行证据映射回需要修改的组件的缺陷诊断机制,在仓库规模上实例化这些原语。在四个软件工程基准上的大量实验表明,HERMES平均比匹配的基线驾驭(harness)高出12.4%。此外,当与强激活和诊断模型配对时,即使使用Qwen3-8B Dev-Primitives,HERMES在所有四个基准上仍保持在同质GPT-5.6 Sol配置的4.5%以内,同时在Terminal-Bench 4.0上将推理成本降低了26.2%,凸显了驾驭设计在软件工程智能体中的重要性。

英文摘要

Large language models (LLMs) equipped with terminal access have demonstrated strong capabilities in automating software engineering tasks. However, existing agents remain brittle on long-horizon workflows, where they must repeatedly reconstruct program state scattered across source files, configurations, tests, dependencies, and runtime behavior, leading to increasingly long interaction histories, context explosion, and semantic drift. Large repositories further complicate the identification of task-relevant components. To address these challenges, we introduce \textbf{Dev-Primitives} (\emph{Development Primitives}), a modular and executable abstraction that transforms repository components from passive software artifacts into active participants in software engineering. Each Dev-Primitive pairs a repository artifact with a resident LLM, which gives the artifact an agent-native interface grounded in its own implementation and dependencies, enabling natural-language reasoning, inter-component communication, and localized self-modification. Building on Dev-Primitives, we propose \textbf{HERMES}, a Harness Engineering framework for software engineeRing via Modular Executable Dev-PrimitiveS, which instantiates these primitives at repository scale through a dependency-aware dynamic activation mechanism and a bug diagnosis mechanism that maps execution evidence back to the components that must be revised. Extensive experiments on four software engineering benchmarks demonstrate that HERMES outperforms matched baseline harnesses by 12.4\% on average. Moreover, when paired with strong activation and diagnosis models, HERMES, even with Qwen3-8B Dev-Primitives, remains within 4.5\% of the homogeneous GPT-5.6 Sol configuration across all four benchmarks, while reducing inference cost by 26.2\% on Terminal-Bench 4.0, highlighting the importance of harness design in software engineering agents.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • The Pennsylvania State University(宾夕法尼亚州立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑