AdviSD:通过目标性多轮自蒸馏学习为前沿大语言模型提供建议
AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation
浏览论文内容
中文总结 AI 辅助
AdviSD通过选择性自蒸馏与强化学习结合,利用反馈修正的评分差异指导训练,使小型顾问有效提升前沿LLM执行器的任务表现。
中文摘要 AI 辅助
一个可训练的小型顾问模型可以使用自然语言建议来引导冻结的语言模型执行器。除了从任务奖励中学习外,顾问还可以利用已完成交互的反馈来改进其建议。然而,一个看似合理的修正并不一定需要改变执行结果,但从这类修正中学习仍可能影响顾问在其他情境中的未来决策。在共享参数的模型中,我们证明了如果这类修正的目标相对于其他修正的目标对有用建议的偏好较弱,则它们会限制学习效果。与从每次修正中学习相比,较少保留这些修正能够提高模型的最终性能。受此启发,我们的方法——顾问自蒸馏(AdviSD),将基于结果的强化学习与从反馈条件的顾问副本中选择性自蒸馏相结合。反思过程提出修正,顾问对相同的记录执行器响应在有和没有其发出的建议两种情况下进行评分,利用评分差异的幅度来选择用于监督的决策。该方法不需要执行器的似然或额外的执行器 rollout。使用Qwen3-8B顾问为Gemini和Claude进行的实验表明,AdviSD在BFCL-v3上比顾问-GRPO高出4.2-6.4个百分点,在EnvScaler上高出3.9-5.1分。训练后的顾问能够泛化到域外任务,并跨不同执行器版本和模型家族进行迁移。AdviSD还优于匹配数量的随机选择,支持了其选择规则的价值。
英文摘要
A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can still affect the advisor's future decisions in other contexts. In a shared-parameter model, we prove that such corrections can limit learning if their targets favor useful advice less strongly than those of other corrections. Keeping them less often than the rest improves the model's eventual performance compared to learning from every correction. Motivated by this, our method, Advisor Self-Distillation (AdviSD), pairs outcome-based reinforcement learning with self-distillation from a feedback-conditioned copy of the advisor selectively. Reflection proposes corrections, and the advisor scores the same recorded executor response with and without its issued advice, using the magnitude of the difference to select decisions for supervision. This approach does not require executor likelihoods or additional executor rollouts. Experiments with Qwen3-8B advisors for Gemini and Claude show that AdviSD outperforms advisor-GRPO by 4.2-6.4 percentage points on BFCL-v3 and by 3.9-5.1 score points on EnvScaler. The trained advisors generalize to out-of-domain tasks and transfer across different executor versions and model families. AdviSD also beats matched-count random selection, supporting the value of its selection rule.
发表机构
- University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。