AI 中文总结
研究LLM性能依赖推理时控制器的问题,提出CALM框架,将其表述为多任务强化学习,通过模块级分解研究训练策略,评估显示该框架能提升推理时工作流程的泛化能力。
AI 中文摘要
大语言模型(LLM)的性能不仅越来越依赖于基础模型,还取决于用于组织推理的推理时控制器。然而,现有的训练后方法通常针对单一固定交互模式进行优化,尽管实际部署依赖于多种控制器,如思维链、自一致性、辩论、规划和验证管道。这导致训练与部署不匹配,并限制了向新工作流程的迁移。我们引入了CALM(控制器感知语言模型),这是一个在训练循环中明确放置控制器的训练后框架。我们将控制器感知训练后方法表述为基于控制器诱导交互协议的多任务强化学习,其中控制器是可重用局部推理模块的组合。这种结构还在回合级GRPO目标下诱导了混合控制器训练的模块级分解,从而能够系统地研究控制器和模块感知训练策略。我们在保留的控制器组合和更广泛的控制器转移上评估了CALM,表明控制器感知训练后方法在超越单控制器优化的推理时工作流程中提高了泛化能力。
英文摘要
Large language model (LLM) performance increasingly depends not only on the base model, but also on the inference-time controller used to organize reasoning. Existing post-training methods, however, typically optimize for a single fixed interaction pattern, despite real deployments relying on diverse controllers such as Chain-of-Thought, self-consistency, debate, planning, and verification pipelines. This creates a training--deployment mismatch and limits transfer to new workflows. We introduce CALM (Controller-Aware Language Models), a post-training framework that explicitly places controllers in the training loop. We formulate controller-aware post-training as multi-task reinforcement learning over controller-induced interaction protocols, where controllers are compositions of reusable local reasoning modules. This structure also induces a module-level decomposition of mixed-controller training under a turn-level GRPO objective, enabling a systematic study of controller and module-aware training strategies. We evaluate CALM on held-out controller compositions and broader controller shifts, showing that controller-aware post-training improves generalization across inference-time workflows beyond single-controller optimization.