发表机构
Sogang University(西江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对RL推理模型后训练中重放贡献难分离的问题,提出GRPO的组级重放控制原语Headroom-Drift Replay,在多推理基准上提升性能并降低Agentic Search场景的时间成本。
AI 中文摘要
基于强化学习(RL)的推理模型后训练正日益受到重复新rollout生成的瓶颈制约,尤其在智能体(Agentic)场景中,环境交互主导了墙钟时间成本。重放可通过复用过往轨迹降低该负担,但现有方法通常将其嵌入包含探索、经验重构或混合策略优化的更大训练流程中,导致难以分离重放自身的贡献。本文提出聚焦问题:仅靠原则性重放选择能达到何种效果?我们引入Headroom-Drift Replay,这是一种用于GRPO的组级重放控制原语,将复用拆分为两个决策:Headroom按剩余学习价值对存储的组排序,Drift则根据与当前策略的兼容性对其进行门控。新的同策略(on-policy)流保持不变,且该方法未添加任何辅助生成或训练机制。在数学推理、多模态推理和Agentic Search基准测试中,这一单一干预措施在Avg Mean@32指标上优于朴素重放,且与更广泛的重放方法相当或超出;在环境交互主导成本的Agentic Search中,它在显著降低墙钟时间的同时提供了相当的质量。
英文摘要
RL-based post-training for reasoning models is increasingly bottlenecked by repeated fresh rollout generation, particularly in agentic settings where environment interaction dominates wall-clock cost. Replay can reduce this burden by reusing past trajectories, but existing methods typically embed it within larger training pipelines involving exploration, experience restructuring, or mixed-policy optimization. This makes replay's own contribution difficult to isolate. We ask a focused question: how far can principled replay selection alone go? We introduce Headroom-Drift Replay, a group-level replay control primitive for GRPO that separates reuse into two decisions. Headroom ranks stored groups by remaining learning value, while Drift gates them by compatibility with the current policy. The fresh on-policy stream remains unchanged, and the method adds no auxiliary generation or training machinery. Across mathematical reasoning, multimodal reasoning, and Agentic Search benchmarks, this single intervention outperforms naive replay and matches or exceeds broader replay methods on Avg Mean@32. In Agentic Search, where environment interaction dominates cost, it delivers comparable quality at materially lower wall-clock time.
Comments51 pages, 25 figures, 17 tables. Accepted at COLM 2026