发表机构
Chongqing University; University of New South Wales(重庆大学; 新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Graph2Env,利用类型化依赖图DepGraph指导仓库环境构建,通过执行反馈持续细化并持久化修复,在200个Python仓库上实现81.0%的构建成功率,优于现有方法。
AI 中文摘要
编码智能体现在越来越依赖执行来验证其解决方案,这使得构建可靠的执行环境成为一项关键的支持能力。然而,仓库环境构建具有挑战性,因为执行需求分散在仓库工件中,并且可能仅在执行过程中才变得明显。现有的基于智能体的方法通过迭代交互来解决这个问题,但关于当前构建状态的信息,包括已发现的需求、已满足和未解决的先决条件及其依赖关系,可能仍然分布在交互历史中。我们提出Graph2Env,一种以DepGraph为中心的基于智能体的方法,DepGraph是一种类型化依赖图,显式表示仓库执行所需的环境需求、它们的依赖关系及其状态。Graph2Env使用DepGraph来指导环境构建,并通过执行反馈持续细化它,同时将成功的修复持久化为可重放的构建过程。然后,将生成的工件应用于全新环境中,以验证所构建的环境可以被重现。我们在从RATBench和EnvBench中提取的200个Python仓库基准上评估Graph2Env,与静态依赖推断基线(pipreqs)、三个专门的环境构建系统(Repo2Run、RAT和SetupX)以及两个通用编码智能体(SWE-agent和Claude Code)进行比较。Graph2Env实现了81.0%的环境构建成功率(EBSR)和59.3%的环境设置成功率(ESSR),分别比最强基线高出9.5和9.0个百分点。
英文摘要
Coding agents now increasingly rely on execution to validate their solutions, making the construction of reliable execution environments a critical enabling capability. However, repository environment construction is challenging because execution requirements are fragmented across repository artifacts and may only become apparent during execution. Existing agent-based approaches address this problem through iterative interaction, but information about the current construction state, including discovered requirements, satisfied and unresolved prerequisites, and their dependencies, can remain distributed across the interaction history. We present Graph2Env, an agent-based approach centered on DepGraph, a typed dependency graph that explicitly represents the environment requirements needed for repository execution, their dependency relations, and their states. Graph2Env uses DepGraph to guide environment construction and continuously refines it with execution feedback, while persisting successful repairs into a replayable construction procedure. The resulting artifacts are then applied in a fresh environment to verify that the constructed environment can be reproduced. We evaluate Graph2Env on a benchmark of 200 Python repositories drawn from RATBench and EnvBench, against a static dependency-inference baseline (pipreqs), three specialized environment-construction systems (Repo2Run, RAT, and SetupX), and two general-purpose coding agents (SWE-agent and Claude Code). Graph2Env achieves an 81.0% Environment Build Success Rate (EBSR) and a 59.3% Environment Setup Success Rate (ESSR), outperforming the strongest baseline by 9.5 and 9.0 percentage points, respectively.