发表机构
Alibaba Cloud Computing(阿里巴巴云计算)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对自主大语言模型后训练中经验重用的决策问题,提出BCIT方法,通过绑定效果上下文、检查条件等减少有害更新,提升模型质量。
AI 中文摘要
大语言模型具备广泛能力,但将其适配到不断变化的领域、工具和需求通常需要反复后训练。自主系统通过提出更新、训练候选模型并利用评估反馈选择后续提议来自动化该过程的部分环节。随着证据积累,一个核心问题浮现:在后续训练改变父模型后,哪些过去的更新证据仍然可用?更新的效果取决于其父模型、数据和训练阶段。将过去的成功视为无上下文的许可会浪费计算资源;若由此产生的子模型被提升,还会破坏后续训练轨迹。我们将此问题形式化为条件经验迁移,并提出边界校准干预迁移(Boundary-Calibrated Intervention Transfer, BCIT)方法,该方法在权重变更训练前授权经验重用。BCIT将观测到的效果与其源上下文绑定,检查适用性条件,否决带有指定硬性冲突的候选,必要时通过有界训练试验获取当前状态证据。完全训练的候选仍需遵循共享采用规则,仅观测到的事件会扩展记忆。在一个适配金融推理、文本转SQL和函数调用的4B模型上,候选更新在评估的上下文间表现出异构的目标和保留效果。在匹配候选、证据和计算资源的情况下,BCIT授权的有害更新更少,且达到的同等预算下最终模型质量高于所评估的替代方案。这些结果支持将经验授权视为自主后训练中的一个独特问题。
英文摘要
Large language models offer broad capabilities, but adapting them to evolving domains, tools, and requirements often entails repeated post-training. Autonomous systems automate parts of this process by proposing updates, training candidates, and using evaluation feedback to select subsequent proposals. As evidence accumulates, a central problem emerges: which past update evidence remains actionable after subsequent training has changed the parent model? An update's effect depends on its parent, data, and training stage. Treating past success as context-free permission can waste compute. If the resulting child is promoted, it can also degrade the subsequent training trajectory. We formulate this problem as conditional experience transfer and introduce Boundary-Calibrated Intervention Transfer (BCIT), a method that authorizes experience reuse before weight-changing training. BCIT binds an observed effect to its source context, checks applicability conditions, vetoes candidates with named hard conflicts, and obtains current-state evidence through a bounded training trial when needed. Fully trained candidates still face a shared adoption rule, and only observed events extend memory. On one 4B model adapted across finance reasoning, text-to-SQL, and function calling, candidate updates exhibit heterogeneous target and retention effects across the evaluated contexts. Under matched candidates, evidence, and compute, BCIT authorizes fewer harmful updates and attains higher equal-budget final-model quality than the evaluated alternatives. These results support treating experience authorization as a distinct problem in autonomous post-training.