发表机构
Nanyang Technological University; Tsinghua University; Tencent; King’s College London(南洋理工大学; 清华大学; 腾讯; 伦敦国王学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
XRepoSkill通过对比成功与失败轨迹提取可迁移规则,验证其跨仓库泛化性,并在SWE-bench Pro和DeepSWE上以三种LLM取得最高问题解决率,较基线提升显著。
AI 中文摘要
软件工程智能体越来越多地使用从先前经验中提炼的可重用技能来解决仓库级问题,然而此类技能往往难以跨仓库迁移。一个核心挑战在于,成功轨迹中出现的行为不一定对成功结果负责:它可能确实有用,可能只是偶然相关,也可能仅仅是模型的习惯性动作。我们提出XRepoSkill,一种基于轨迹的学习可迁移技能的方法。我们将技能表示为一系列规则,每条规则指定在问题解决过程中采取什么行动以及何时采取。XRepoSkill首先对比同一智能体在相同问题上的成功与失败轨迹,并从其执行路径分叉处推导候选规则。每条规则配有一个可执行的谓词,使其规定行为能够在其他轨迹上被系统评估。规则根据其与成功解决问题的关联性进行验证,仅当其规定行为在多个仓库中重复出现时才被保留;同一行为的仓库特定变体随后被整合为可迁移规则。对于新问题,XRepoSkill选择相关规则来指导智能体。我们从官方SWE-bench Verified排行榜上公开的轨迹中学习技能,并使用来自三个不同供应商的后端大语言模型在SWE-bench Pro和DeepSWE上进行评估;所有评估仓库均未出现在技能学习轨迹池中。与三种近期技能学习方法相比,XRepoSkill在所有六种基准测试与LLM组合中取得了最高的问题解决率。特别是在具有挑战性的长时程DeepSWE基准上,XRepoSkill将问题解决率比未使用学习技能的同一智能体提高了10.3个百分点,比最强的技能学习基线提高了5.0个百分点。
英文摘要
Software engineering agents increasingly use reusable skills distilled from prior experience to resolve repository-level issues, yet such skills often fail to transfer across repositories. A central challenge is that a behavior appearing in a successful trajectory is not necessarily responsible for the successful outcome: it may be genuinely useful, merely incidental, or simply a recurring habit of the model. We introduce XRepoSkill, a trajectory-based approach for learning transferable skills. We represent a skill as a collection of rules, each specifying what action to take and when to take it during issue resolution. XRepoSkill first contrasts successful and failed trajectories of the same agent on the same issue and derives candidate rules from where their execution paths diverge. Each rule is paired with an executable predicate that enables its prescribed behavior to be evaluated systematically on other trajectories. A rule is verified based on its association with successful issue resolution and retained only when its prescribed behavior recurs across multiple repositories; repository-specific variants of the same behavior are then consolidated into transferable rules. For a new issue, XRepoSkill selects relevant rules to guide the agent. We learn skills from publicly released trajectories on the official SWE-bench Verified leaderboard and evaluate them on SWE-bench Pro and DeepSWE using three backbone LLMs from different vendors; none of the evaluation repositories appears in the skill-learning trajectory pool. Against three recent skill learning methods, XRepoSkill achieves the highest issue resolution rate in all six benchmark--LLM combinations. In particular, on the challenging long-horizon DeepSWE benchmark, XRepoSkill improves issue resolution by 10.3 percentage points over the same agent without learned skills and by 5.0 points over the strongest skill-learning baseline.