AI 中文总结
PG-SFT通过利用智能体轨迹的回合级信息增益调整监督强度,在离线智能体微调中平衡能力获取与保留,显著减少分布漂移和广泛能力退化,仅以轻微目标任务性能下降为代价。
AI 中文摘要
在离线智能体轨迹上进行监督微调(SFT)是训练专用工具使用智能体的标准方法,但强制模型逐词模仿推理和行动可能会损害基础模型的其他能力(如通用推理、工具调用、代码生成)。在本工作中,我们专注于研究如何在智能体轨迹SFT过程中更好地平衡获取新能力与保留已有能力之间的权衡。通过在我们设置中比较几个基线,标准SFT提高了目标基准的性能,同时降低了几个非目标基准的分数;与此同时,简单地使用KL惩罚约束分布漂移或限制更新幅度并未避免这种退化趋势。受近期逐词自适应学习目标的启发,本工作提出特权引导SFT(PG-SFT),利用智能体轨迹的回合级信息增益作为调整监督强度的指标。PG-SFT在评估基准上产生了更有利的观察权衡,显著减少了分布漂移和广泛能力退化,代价是目标任务性能的轻微下降。我们的发现表明,平衡获取-保留权衡不仅取决于模型是否锚定于其基础行为,还取决于监督应在何处以及以何种强度偏离该行为。
英文摘要
Supervised fine-tuning (SFT) on offline agent trajectories is the standard approach for training specialized tool-using agents, but forcing models to imitate reasoning and actions token by token may harm other capabilities (e.g., general reasoning, tool calling, code generation) of the base model. In this work, we focus on studying \emph{how to better balance the trade-off between acquiring new capabilities and preserving existing ones during agent trace SFT}. By comparing several baselines in our setup, standard SFT improves the target benchmark while lowering several non-target benchmark scores; meanwhile, simply constraining distributional drift using KL penalty or limiting the update magnitude did not avoid this regression trend. Motivated by recent token-wise adaptive learning objectives, this work proposes \textbf{Privilege-Guided SFT (PG-SFT)} to leverage turn-level information gain of agent trajectories as an indicator to adjust supervision strength. PG-SFT yields a more favorable observed trade-off on the evaluated benchmarks, substantially reducing distributional drift and broad capability degradation at the cost of slight degradation in target-task performance. Our findings suggest that balancing the acquisition--retention trade-off depends not only on whether the model is anchored to its base behavior, but also on where and how strongly supervision should depart from that behavior.}