AI 中文总结
针对大型推理模型过度思考导致的高推理成本问题,提出奖励协调压缩框架ReCo,通过奖励自适应缓存压缩等组件,减少生成Token并降低延迟,同时保持准确率。
AI 中文摘要
大型推理模型(LRMs)通过长思维链(CoT)推理在复杂任务上表现出色,但其冗长的中间步骤会导致严重的过度思考,从而增加推理成本。KV缓存压缩是一种常见的解决方案,但现有的面向推理的方法会在整个轨迹中应用统一策略,仅通过从缓存中移除的内容来判断压缩效果。两个观察结果指向了另一种思路:第一,推理状态对上下文丢失的容忍度沿轨迹变化,过程奖励可追踪这种变化——在高奖励步骤删除Token比随机删除相同预算更能保持准确率;第二,生成侧的压缩并非无代价,因为更小的缓存会导致模型生成更多Token,部分抵消了节省的成本。这些共同促使我们在单一过程奖励下协调两侧。我们提出ReCo(奖励协调压缩),这是一个逐步框架,其中轻量级过程奖励估计器对每个已完成步骤进行评分,并驱动三个组件:(1)奖励自适应KV缓存压缩,在高奖励步骤更大力地收缩保留的缓存,在低奖励步骤则收缩力度较小;(2)对反思Token的奖励带惩罚,以抑制冗余生成;(3)基于置信度的提前停止,当推理可靠时触发。在三个推理模型和六个基准上,与完整思维链(Full CoT)相比,ReCo减少了37%-65%的生成Token,端到端延迟降低了2.08倍-2.35倍,同时在很大程度上保持了准确率。
英文摘要
Large Reasoning Models (LRMs) excel on complex tasks through long chain-of-thought (CoT) reasoning, but their lengthy intermediate steps cause severe overthinking that inflates inference cost. KV-cache compression is a common solution, yet existing reasoning-oriented methods apply a uniform policy across the trajectory and judge compression only by what it removes from the cache. Two observations point the other way. First, a reasoning state's tolerance to context loss varies along the trajectory, and process reward tracks it: deleting tokens at high-reward steps preserves accuracy far better than deleting the same budget at random. Second, compression is not free on the generation side, since a smaller cache leads the model to generate more tokens, partly canceling the saving. Together these motivate coordinating both sides under a single process reward. We propose ReCo (Reward-Coordinated Compression), a step-wise framework in which a lightweight process-reward estimator scores each completed step and drives three components: (1) reward-adaptive KV-cache compression that shrinks the retained cache harder at high-reward steps and less at low-reward ones, (2) a reward-banded penalty on reflection tokens that curbs redundant generation, and (3) confidence-based early stopping that triggers when the reasoning is reliable. Across three reasoning models and six benchmarks, ReCo reduces generated tokens by 37%-65% and end-to-end latency by 2.08x-2.35x over Full CoT, all while largely preserving accuracy.
CommentsWork in progress, revisions ongoing