发表机构
Southern University of Science and Technology; The Hong Kong University of Science and Technology; Shanghai Artificial Intelligence Laboratory; Tongji University(南方科技大学; 香港科技大学; 上海人工智能实验室; 同济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MetaCtrl提出轻量级元认知控制器,通过强化学习动态调节冻结推理者的推理过程,在七个基准上提升准确性并缩短推理长度,且可迁移至新模型。
AI 中文摘要
大型推理模型通过在回答前分配额外计算来提升在难题上的性能,但更长的推理并不总是带来更好的结果,反而可能在简单问题上引入大量冗余推理。相反,激进地缩短推理会降低在难题上的表现。因此,有效推理需要基于推理者的能力和不断演化的解决方案状态,动态决定何时额外计算是有用的。现有方法通常依赖预定义的预算或干预规则、重新训练目标推理者,或需要额外的监督。我们提出MetaCtrl,一种轻量级控制器,无需预定义令牌预算或推理者重训练即可自适应地调节冻结的推理者。我们将推理调节形式化为一个序列元认知控制问题:MetaCtrl观察不断演化的推理轨迹,并决定是继续、简化、跳过冗余步骤还是结束推理。它直接通过强化学习训练,使用优先考虑正确性并在正确解决方案中偏好较短轨迹的奖励,既不需要监督干预轨迹,也不需要特定问题的预算。在涵盖数学、科学和代码的七个基准上,MetaCtrl持续提升大型推理模型的准确性,同时减少其推理长度。在DeepSeek-R1-Distill-Qwen-7B上,它平均准确率提升4.7个百分点,同时生成长度减少53.3%。无需进一步训练,同一控制器可迁移至未见过的推理者(如Qwen3-14B),平均准确率提升2.9个百分点,生成长度减少50.3%。这些结果确立了MetaCtrl作为即插即用控制器,在显著减少推理时生成的同时提升推理准确性。代码可在该https URL获取。
英文摘要
Large reasoning models improve performance on challenging problems by allocating additional computation before answering, but longer reasoning does not always lead to better results and can introduce substantial redundant reasoning on simple problems. Conversely, aggressively shortening reasoning can degrade performance on difficult ones. Effective reasoning therefore requires dynamically deciding when additional computation is useful based on the reasoner's capabilities and evolving solution state. Existing approaches often rely on predefined budgets or intervention rules, retrain the target reasoner, or require additional supervision. We introduce MetaCtrl, a lightweight controller that adaptively regulates a frozen reasoner without predefined token budgets or reasoner retraining. We formulate reasoning regulation as a sequential metacognitive control problem: MetaCtrl observes the evolving reasoning trace and decides whether to continue, simplify, skip redundant steps, or conclude reasoning. It is trained directly with reinforcement learning using a reward that prioritizes correctness while favoring shorter trajectories among correct solutions, requiring neither supervised intervention trajectories nor problem-specific budgets. Across seven benchmarks spanning mathematics, science, and code, MetaCtrl consistently improves the accuracy of LRMs while reducing their reasoning length. On DeepSeek-R1-Distill-Qwen-7B, it improves average accuracy by 4.7 points while reducing generation length by 53.3%. Without further training, the same controller transfers to an unseen reasoner (e.g., Qwen3-14B), improving average accuracy by 2.9 points and reducing generation length by 50.3%. These results establish MetaCtrl as a plug-and-play controller for improving reasoning accuracy while substantially reducing inference-time generation. The code is available at https://github.com/binbin2xs/MetaCtrl.