arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00332cs.LGcs.AI

最薄弱环节:基于最坏情况约束强化学习的大语言模型推理蒸馏

The Weakest Link: Distilling LLM Reasoning with Worst-Case Constrained Reinforcement Learning

Matthieu Zimmer, Xiaotong Ji, Tu Nguyen, Haitham Bou-Ammar

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM推理蒸馏中奖励黑客与平均化缺陷,提出最坏情况约束RL方法,通过无增强MDP变换保留硬约束,显著扩展准确率-保真度帕累托前沿,实现最高严格推理成功率。

中文摘要 AI 辅助

将大语言模型(LLM)的推理能力蒸馏到较小的学生模型中是高效部署的核心挑战。现有方法面临一个基本矛盾:纯粹优化可验证的任务奖励(例如通过GRPO)会导致奖励黑客行为,即学生通过有缺陷的中间逻辑得出正确的最终答案;而使用针对教师的软散度惩罚进行正则化(例如基于KL的蒸馏)则会稀释任务性能,并且关键的是,允许学生用某一步骤的高教师一致性来弥补另一步骤的严重逻辑违规。我们认为这种平均化从根本上与推理的本质不符:思维链的有效性取决于其最薄弱的环节。受此观察启发,我们将推理蒸馏形式化为一个受约束的强化学习问题,其中任务奖励在轨迹每个前缀上的教师对数似然最坏情况约束下最大化。为避免对偶拉格朗日求解器的过高成本以及Saute等状态增强方法在测试时对教师的依赖,我们推导出一个无增强的受约束MDP,其奖励变换保留了硬约束语义,允许将策略梯度低方差分解为单步项和长期项,并在惩罚极限下几乎必然满足最坏情况约束。通过在数学推理和代码生成任务上的大量实验,我们证明该方法显著扩展了准确率-保真度帕累托前沿。通过匹配纯RL的高最终答案正确率并大幅减少教师约束违规,我们最终在所有评估设置中实现了最高的严格推理成功率。

英文摘要

Distilling the reasoning capabilities of large language models (LLMs) into smaller students is a central challenge for efficient deployment. Current approaches face a fundamental tension: optimizing purely for verifiable task rewards (e.g., via GRPO) leads to reward hacking, where students arrive at correct final answers through flawed intermediate logic, while regularizing with soft divergence penalties against a teacher (e.g., KL-based distillation) dilutes task performance and, critically, allows the student to compensate for severe logical violations at one step with high teacher agreement at others. We argue that this averaging is fundamentally misaligned with the nature of reasoning: a chain-of-thought is only as valid as its weakest link. Motivated by this observation, we formulate reasoning distillation as a constrained reinforcement learning problem in which the task reward is maximized subject to a worst-case constraint on the teacher log-likelihood along every prefix of the trajectory. To avoid the prohibitive cost of dual Lagrangian solvers and the test-time teacher dependence of state-augmented methods such as Saute, we derive an unaugmented constrained MDP whose reward transformation preserves the hard-constraint semantics, admits a low-variance policy gradient decomposition into single-step and long-term terms, and provably satisfies the worst-case constraint almost surely in the penalty limit. Through extensive experiments on mathematical reasoning and code generation tasks, we demonstrate that our method significantly expands the accuracy-fidelity Pareto front. By matching the high Final Answer Correctness of pure RL and drastically reducing teacher constraint violations, we ultimately achieve the highest rigorous Reasoning Success Rate across all evaluated settings.

发表机构

  • Huawei Heisenberg Research Center(华为海森堡研究中心)
  • UCL Centre for AI(伦敦大学学院人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

↑