CounterRoute:通过分层反事实信用分配实现自路由推理
CounterRoute: Self-Routed Reasoning via Hierarchical Counterfactual Credit Assignment
AI总结:
CounterRoute提出在线强化学习框架,联合学习路由与模式条件响应,通过反事实信用分配和课程训练,在九个基准上提升准确率并大幅减少生成令牌,且泛化良好。
AI中文摘要:
具备推理能力的语言模型在直接回答即可满足需求时,往往会产生冗长的思维链,从而浪费推理计算资源。许多双模式模型将这一选择留给用户。自动化这一选择具有挑战性,因为路由目标随策略演化而改变,初始模式偏好会破坏探索的稳定性,且序列级目标将路由与响应学习纠缠在一起。我们提出CounterRoute,一个在线强化学习框架,直接从原生双模式检查点出发,在单一共享策略中联合学习路由和模式条件响应,无需特定方法的SFT预热。配对当前策略的反事实轨迹仅将跨模式信用分配给路由令牌,而模式内GRPO训练响应令牌。配对到自路由的课程通过强制两种模式的轨迹稳定早期训练,随后增加自路由更新以提升自主路由能力。在九个基准测试中,CounterRoute在准确性和效率之间取得了比启发式和学习的自适应路由方法更好的平衡。相对于始终思考的检查点,它在提高宏平均准确率的同时,将Qwen3-8B的平均生成令牌数减少51%,Qwen3-14B减少41%。在直接回答能力较强的指令遵循和常识基准上,思考率低至1%,而响应质量有所提升。尽管仅基于数学和指令遵循数据进行训练,其路由行为和响应质量仍能泛化到保留的编码、科学、知识和常识基准。
英文摘要:
Reasoning-capable language models often produce long chains of thought when direct answers suffice, wasting inference compute. Dual-mode models offer both thinking and direct-answer modes, but typically leave mode selection to users. Automating this selection while improving responses under both modes is challenging because routing targets evolve with the policy, strong initial mode preferences can destabilize exploration, and sequence-level objectives entangle routing with response learning. We introduce CounterRoute, an online reinforcement-learning framework that jointly learns routing and mode-conditioned responses in one shared policy directly from a native dual-mode checkpoint, without method-specific SFT warm-up. Paired current-policy counterfactual rollouts estimate cross-mode routing credit, which is assigned only to the routing token, while within-mode GRPO trains response tokens. A paired-to-self-routed curriculum stabilizes early training with forced coverage of both modes before progressively transferring training to autonomous routing. Across nine benchmarks, CounterRoute better balances accuracy and efficiency than heuristic and learned adaptive-routing methods, improving macro-average accuracy at both Qwen3 scales while cutting mean generated tokens by 51% and 41%, respectively, versus always-thinking checkpoints. On instruction-following and commonsense benchmarks where direct answering is strong, think rates fall as low as 1% while response quality improves. Despite training only on math and synthetic instruction-following data, its routing and response quality generalize to held-out tasks in coding, science, knowledge, and commonsense.