arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27449cs.SEcs.AIcs.CL

SWE-Prime:更少轨迹,更优性能

SWE-Prime: Fewer Trajectories, Better Performance

Dewu Zheng, Ruizhe Ye, Yanlin Wang, Yang Ye, Hongyu Zhang, Ensheng Shi, Xilin Liu, Yuchi Ma, Jianxing Yu, Zibin Zheng

首次发表
浏览论文内容

中文总结 AI 辅助

SWE-Prime是一种多粒度两阶段SFT数据选择方法,在SWE-Bench基准上仅用10%轨迹训练,就实现了优于全数据集训练的软件问题解决性能,相对提升最高达24.2%。

中文摘要 AI 辅助

为提升大语言模型解决现实世界软件问题的能力,现有研究多聚焦于构建大规模智能体轨迹数据集,并在成功轨迹上执行监督微调(SFT)。然而任务成功并不等同于高质量监督:成功轨迹可能仍包含无效、冗余或有风险的步骤,直接将此类轨迹用于SFT会引入噪声监督,促使模型模仿不良的问题解决行为。因此,本文提出SWE-Prime,一种多粒度、两阶段的SFT数据选择方法,逐步在轨迹和片段层级过滤训练数据。具体而言,第一阶段基于过程质量、结果质量和数据代表性进行轨迹层级筛选,选取高质量且具代表性的成功轨迹子集;第二阶段将连续步骤分组为语义片段,基于各片段对最终解决方案的贡献、可学习性及潜在风险进行片段层级选择。在SFT过程中,所有片段均保留在序列中以维持上下文,仅选定的片段参与损失计算。在SWE-Bench Pro和SWE-Bench Verified上的实验表明,使用SWE-Prime选定的10%轨迹子集进行训练,其性能优于使用全部已解决数据集训练,分别实现了最高12.2%和24.2%的相对性能提升。

英文摘要

To improve large language models' ability to resolve real-world software issues, prior work has focused on constructing large-scale agent trajectory datasets and performing supervised fine-tuning (SFT) on successful trajectories. However, task success does not guarantee high-quality supervision: successful trajectories may still contain ineffective, redundant, or risky steps. Directly using such trajectories for SFT can introduce noisy supervision and encourage models to imitate undesirable problem-solving behaviors. Therefore, we propose SWE-Prime, a multi-granularity, two-stage SFT data selection method that progressively filters training data at the trajectory and segment levels. Specifically, the first stage performs trajectory-level screening based on process quality, result quality, and data representativeness, selecting a high-quality and representative subset of successful trajectories. The second stage performs segment-level selection by grouping consecutive steps into semantic segments and assessing each segment based on its contribution to the final solution, learnability, and potential risks. During SFT, all segments remain in the sequence to preserve context, while only selected segments contribute to the loss computation. Experiments on SWE-Bench Pro and SWE-Bench Verified show that training on the 10% trajectory subset selected by SWE-Prime outperforms training on the full resolved dataset, yielding relative performance gains of up to 12.2% and 24.2%, respectively.

发表机构

  • Sun Yat-sen University(中山大学)
  • Huawei Cloud Computing Technologies Co., Ltd.(华为云计算技术有限公司)
  • Chongqing University(重庆大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑