arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17205cs.AIcs.SE

代码智能体LoRA微调的轨迹数据管理系统评估

A Systematic Evaluation of Trajectory Data Curation for LoRA Fine-Tuning of Code Agents

Yunze Han

首次发表
浏览论文内容

中文总结 AI 辅助

研究代码智能体LoRA微调中轨迹数据管理,提出效率和风格两轴质量评分框架,经16个控制实验评估,揭示规模依赖的质量-数量权衡及错误重试率是主要子维度,为代码智能体监督微调提供可行手段与评估协议。

中文摘要 AI 辅助

对开放权重语言模型在专家智能体轨迹上进行监督微调已成为构建高性能代码智能体的重要方法,而轨迹质量和数量如何共同影响模型性能是一个未充分探索的核心问题。本文对Qwen2.5-Coder-7B-Instruct在SWE轨迹数据集上进行LoRA微调的轨迹数据过滤进行了系统实证研究。提出了效率和风格两轴质量评分框架,并通过16个控制实验进行评估。由于7B规模模型的SWE基准解析率接近零,采用交叉熵损失作为主要指标,并通过首次动作生成进行验证。结果揭示了规模依赖的质量-数量权衡,消融分析还确定了错误重试率是主要子维度。这些发现确立了轨迹级质量评分是代码智能体监督微调的可行但对规模敏感的手段,并提供了一个代理验证的评估协议。

英文摘要

Supervised fine-tuning (SFT) of open-weight LLMs on expert agent trajectories has emerged as a prominent approach to building capable code agents without reliance on proprietary models. A central yet underexplored question is how trajectory quality and quantity jointly shape model performance. We present a systematic empirical study of trajectory data filtering for LoRA fine-tuning of Qwen2.5-Coder-7B-Instruct on the SWE-trajectory dataset (67,074 trajectories, of which 32,161 are resolved). We propose a two-axis quality scoring framework -- Efficiency and Style -- and evaluate it through 16 controlled experiments spanning strategy, scale, and ablation analyses. Since 7B-scale models attain near-zero SWE-bench resolve rates, we adopt cross-entropy (CE) loss on held-out trajectories as the primary metric, validated via first-action generation: CE loss and ROUGE-L are perfectly rank-correlated (Spearman $ρ$ = -1.00), with limited-sample evidence supporting but not conclusively establishing this proxy. Our results reveal a scale-dependent quality-quantity trade-off: at small scales, doubling the dataset (500 to 1,000) yields ~12.7% CE-loss reduction whereas the TopQ-Random gap stays <1% (Mann-Whitney p > 0.10); at 2,000 trajectories this same gap widens to 3.6% (p = 0.016). Ablation further identifies error-retry rate as the dominant sub-dimension, performing comparably to the full composite ($Δ$ < 0.2%). Together, these findings establish trajectory-level quality scoring as a viable but scale-sensitive lever for code-agent SFT and offer a proxy-validated evaluation protocol for the regime where end-to-end resolve rate is statistically infeasible.

补充信息

↑