Stable-MM-R1:通过熵引导分层锚定多模态推理动态
Stable-MM-R1: Anchoring Multimodal Reasoning Dynamics via Entropy-Guided Stratification
浏览论文内容
中文总结 AI 辅助
针对强化学习训练不稳定和熵崩溃问题,提出以数据为中心的PAQM和HSR框架,通过熵引导分层稳定RL微调,在复杂推理任务上超越强基线。
中文摘要 AI 辅助
虽然强化学习(RL)能有效激励大型语言模型中的推理能力,但当前的流程受到训练不稳定和熵快速崩溃的阻碍。这些限制通常源于标准采样过程中的“回滚沉默”和低质量梯度信号。在本工作中,我们提出了一个稳健的、以数据为中心的框架来稳定RL训练。我们首先引入潜在感知查询挖掘(PAQM),该机制动态过滤数据,聚焦于“蒸馏区”——即具有高能力激发潜力的样本。此外,我们提出了混合分层回放(HSR),这是一种新颖的机制,通过基于路径熵(一种回滚级别的置信度代理)和结果奖励对回放进行分层来重构批次。在每个优化步骤中,HSR重用当前策略的“稳定性锚点”和“硬负样本”来构建高对比度的优化组,然后在下一步之前清空其缓冲区。这种方法在有限计算下缓解了熵崩溃,同时提高了学习信号的利用率。我们的方法在复杂推理任务上优于强基线,为稳定高效的RL微调提供了一种原则性解决方案。
英文摘要
While Reinforcement Learning (RL) effectively incentivizes reasoning in Large Language Models, current pipelines are hindered by training instability and rapid entropy collapse. These limitations often stem from "Rollout Silencing" and low-quality gradient signals in standard sampling procedures. In this work, we propose a robust, data-centric framework to stabilize RL training. We first introduce Potential-Aware Query Mining (PAQM), which filters data dynamically to focus on the "Distillation Zone"---samples with high potential for capability elicitation. Furthermore, we present Hybrid Stratified Replay (HSR), a novel mechanism that restructures batches by stratifying rollouts based on Path Entropy, a rollout-level confidence proxy, and outcome reward. Within each optimization step, HSR reuses current-policy "Stability Anchors" and "Hard Negatives" to construct high-contrast optimization groups, then clears its buffers before the next step. This approach mitigates entropy collapse while improving the utilization of learning signals under limited compute. Our method outperforms strong baselines on complex reasoning tasks, offering a principled solution for stable and efficient RL fine-tuning.
发表机构
- Columbia University(哥伦比亚大学)
- CUHK(香港中文大学)
- THU(清华大学)
- McGill University(麦吉尔大学)
机构由 AI 辅助整理,请以论文原文为准。