arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于非对称强化学习与自蒸馏的指令条件探索

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy

Jim Dilkes, Vahid Yazdanpanah, Sebastian Stein

arXiv 2608.02087首次发表:更新:

发表机构

University of Southampton(南安普顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对LLM强化学习中的探索挑战,提出指令条件探索(ICE)方法,结合非对称RL/SD训练目标,使Qwen3-1.7B数学推理性能提升5.0%且长上下文下仍有效。

AI 中文摘要

利用强化学习(RL)对大语言模型(LLMs)进行后训练已成为提升模型能力的重要工具,但LLMs的动作空间结构带来了与经典RL不同的挑战,对诱导探索具有重要影响。需要新方法来利用预训练LLMs的广泛知识和灵活性,在训练时刻意生成多样化经验。我们提出指令条件探索(ICE),在训练时给任务提示补充多种不同指令中的一种,增加尝试行为的覆盖范围。为实现ICE,我们提出非对称RL/SD,一种结合强化学习与自蒸馏的训练目标,将探索到的行为迁移到无指令的测试时策略。采用非对称RL/SD目标的ICE,使Qwen3-1.7B在数学推理任务上,4K响应长度的保留pass@1性能相比使用DAPO训练提升了5.0%,且在更长的8K上下文下该提升依然存在。

英文摘要

Post-training Large Language Models (LLMs) with Reinforcement Learning (RL) has become an important tool for improving model capabilities, but the LLM action-space structure introduces challenges distinct from classical RL, with implications for inducing exploration. New methods are required that leverage the broad knowledge and flexibility of pre-trained LLMs to deliberately generate diverse experience at training time. We propose Instruction-Conditioned Exploration (ICE), which appends one of a small fixed set of instructions to task prompts during training, using the same set for every problem, increasing the coverage of behaviours attempted. To facilitate ICE, we combine RL on the instruction-conditioned policy with self-distillation of its correct rollouts into the unconditioned test-time policy. ICE with this objective improves Qwen3-1.7B held-out pass@1 performance at 4K response length on mathematical reasoning tasks by $5.0\%$ relative to training with DAPO, with improvement persisting at a longer 8K context. The improvement does not appear for Qwen3-4B at 4K, where the instructions do not expand base-model coverage.

CommentsSubmitted to ACL Rolling Review (ARR) May 2026 cycle. OpenReview submission record at https://openreview.net/forum?id=PV945lekMa

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑