从轨迹到前缀:通过重放前缀和在线延续复用教师轨迹
From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation
浏览论文内容
中文总结 AI 辅助
研究针对小语言模型蒸馏低效问题,提出Prefix-GRPO强化学习框架,将教师轨迹分解,通过重放前缀恢复中间状态并在线延续,统一前缀与延续学习,实验证明其优于蒸馏和标准RL基线,凸显前缀令牌优化的重要性。
中文摘要 AI 辅助
小语言模型是交互式智能体的有吸引力的主干,但直接从强大的教师轨迹进行蒸馏往往会将丰富的多轮行为转化为一次性模仿目标。这在长期环境中效率低下,早期决策会影响后期状态和奖励。我们提出Prefix-GRPO,一种强化学习框架,将教师轨迹分解为重放对齐的前缀查询和在线延续。每个前缀在环境中重放以恢复有效的中间状态,然后学生继续在线交互并接收任务奖励。与仅响应的GRPO不同,Prefix-GRPO还对重放前缀内的历史辅助令牌应用裁剪策略更新,使用策略蒸馏的SFT检查点估计其旧对数概率。这在同一策略优化形式中统一了前缀学习和延续学习。在TextCraft、BabyAI和ALFWorld上的实验表明,Prefix-GRPO比蒸馏和标准RL基线改进了小模型智能体,而消融实验表明,没有显式前缀令牌优化,仅重放是不够的。实现和再现脚本可在该https URL获得。
英文摘要
Small language models are attractive backbones for interactive agents, but direct distillation from strong teacher trajectories often turns rich multi-turn behavior into one-shot imitation targets. This is inefficient in long-horizon environments, where early decisions shape later states and rewards. We propose Prefix-GRPO, a reinforcement learning framework that decomposes teacher trajectories into replay-aligned prefix queries and online continuations. Each prefix is replayed in the environment to recover a valid intermediate state, after which the student continues online interaction and receives task reward. Unlike response-only GRPO, Prefix-GRPO also applies clipped policy updates to historical assistant tokens inside the replayed prefix, using a policy-distilled SFT checkpoint to estimate their old log-probabilities. This unifies prefix learning and continuation learning within the same policy-optimization form. Experiments on TextCraft, BabyAI, and ALFWorld show that Prefix-GRPO improves small-model agents over distillation and standard RL baselines, while ablations show that replay alone is insufficient without explicit prefix-token optimization. The implementation and reproduction scripts are available at https://github.com/HappynessI/Prefix_GRPO.
发表机构
- Tianjin University(天津大学)
- Peking University(北京大学)
- Tianjin University of Technology(天津工业大学)
机构由 AI 辅助整理,请以论文原文为准。