训练期间模型行为的可扩展归因与控制
Scalable Attribution and Control of Model Behavior During Training
浏览论文内容
中文总结 AI 辅助
本文提出基于互信息的BGU指标及BS-Ghost算法,实现在训练中快速归因样本贡献并预测干预效果,从而引导模型行为,开销仅增8%。
中文摘要 AI 辅助
在训练期间归因和控制模型行为,需要在下次更新之前足够快地识别每个样本的贡献。然而,同一训练批次中的样本可能产生相似的行为变化,使得它们的个体贡献难以区分。我们通过互信息来解决这种模糊性,通过量化组合行为变化对每个样本贡献的揭示程度来考虑批次内的干扰。我们证明该互信息是行为梯度唯一性(BGU)的对数函数。BGU为信息度量提供了几何解释。我们的批空间幽灵(BS-Ghost)算法通过批空间中的共享计算,使这些分数在训练循环内变得实用,无需存储模型大小的样本梯度。在完整的1,000样本Qwen2.5-7B-Instruct工作负载上,我们的BS-Ghost实现为5.5分钟的普通训练增加了27秒(8.0%)。移除和重新训练实验表明,BGU能够识别因果塑造最终行为的数据。在每个训练步骤中,带符号的信息识别哪些样本增强或削弱目标行为,从而解释行为在训练期间如何发展。带符号的信息还支持训练期间的廉价干预:它预测改变样本权重将如何影响下一次更新中的行为。然后我们利用这些预测来选择权重,将行为引导向期望的目标。这使得我们的框架成为训练流程和过程的可扩展监督与验证的实用基础,帮助评估者评估模型对齐,理解其在训练期间如何发展,并指导塑造持续学习的干预措施。
英文摘要
Attributing and controlling model behavior during training requires identifying each example's contribution quickly enough to act before the next update. However, examples in the same training batch can produce similar behavioral changes, making their individual contributions difficult to distinguish. We address this ambiguity through mutual information, accounting for interference within the batch by quantifying how much the combined behavioral change reveals about each example's contribution. We show that this mutual information is a logarithmic function of Behavioral Gradient Uniqueness (BGU). BGU gives the information measure its geometric interpretation. Our Batch-Space Ghost (BS-Ghost) algorithm makes these scores practical inside the training loop through shared computation in batch space, without storing model-sized example gradients. On a complete 1,000-example Qwen2.5-7B-Instruct workload, our BS-Ghost implementation adds 27 seconds (8.0%) to 5.5 minutes of ordinary training. Removal and retraining demonstrate that BGU identifies data that causally shapes final behavior. At each training step, signed information identifies which examples strengthen or weaken the target behavior, explaining how behavior develops during training. Signed information also enables cheap intervention during training: it predicts how changing example weights will affect behavior in the next update. We then use these predictions to choose weights that steer behavior toward a desired target. This makes our framework a practical foundation for scalable oversight and verification of training pipelines and processes, helping evaluators assess model alignment, understand how it develops during training, and guide interventions that shape ongoing learning.
发表机构
- Rice University(莱斯大学)
机构由 AI 辅助整理,请以论文原文为准。