发表机构
Shenzhen Technology University; Shenzhen University; Shandong University; Huazhong University of Science and Technology; Institute of Software, Chinese Academy of Sciences(深圳理工大学; 深圳大学; 山东大学; 华中科技大学; 中国科学院软件研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EnvPilot通过轨迹派生记忆和上下文感知检索复用历史经验,在AES-Bench基准上以75%的Pass@1实现软件环境搭建的新SOTA,并降低推理成本。
AI 中文摘要
环境搭建是软件工程中一项关键但复杂的任务,高度依赖专家知识。现有的自动化环境搭建方法缺乏从过往执行轨迹中积累经验并随时间演进的能力。因此,它们的性能受限,因为它们常常进行冗余探索、忽略有用的历史解决方案,并且难以在不同软件生态系统中泛化。我们提出了EnvPilot的系统设计与实证验证,EnvPilot是一个经验增强的智能体,它将基于轨迹的经验复用操作化,用于软件环境搭建。EnvPilot维护一个可扩展的轨迹派生记忆(TDM),初始包含667条高质量经验。它系统性地将历史执行轨迹中的隐性知识转化为结构化经验,并通过上下文感知检索机制在任务执行期间检索最相关的指导。这使得EnvPilot能够组合多种经过验证的搭建策略,提供比仅依赖静态项目文件或网络检索的方法更精确、更详细的指导。为评估EnvPilot,我们构建了AES-Bench,一个包含112个真实世界GitHub实例、覆盖9种编程语言的多语言基准。实验表明,EnvPilot取得了新的最先进(SOTA)成果,Pass@1成功率达75.00%,同时降低了推理成本。我们的实证研究表明,结构化经验表示和上下文感知检索机制都是不可或缺的。
英文摘要
Environment Setup is a critical yet complex task in software engineering that relies heavily on expert knowledge. Existing automated environment setup methods lack the ability to accumulate experience from past execution trajectories and to evolve over time. As a result, their performance is limited because they often perform redundant exploration, ignore useful past solutions, and fail to generalize across diverse software ecosystems. We present the systematic design and empirical validation of EnvPilot, an experience-augmented agent that operationalizes trajectory-derived experience reuse for software environment setup. EnvPilot maintains an expandable Trajectory-Derived Memory (TDM), initialized with 667 high-quality experiences. It systematically transforms implicit knowledge from historical execution trajectories into structured experience and retrieves the most relevant guidance during task execution through the Context-aware Retrieval mechanism. This enables EnvPilot to combine multiple validated setup strategies, providing more precise and detailed guidance than methods that rely solely on static project files or web retrieval. To evaluate EnvPilot, we construct AES-Bench, a multilingual benchmark of 112 real-world GitHub instances across 9 programming languages. Experiments show that EnvPilot achieves a new state-of-the-art (SOTA) with a 75.00% Pass@1 success rate while reducing reasoning costs. Our empirical study shows that both the structured experience representation and the Context-aware Retrieval mechanism are essential.
Comments49 pages, 8 figures. Accepted for publication in ACM Transactions on Software Engineering and Methodology (TOSEM)