AI 中文总结
研究针对紧凑模型训练成本高问题,提出MADA-RL框架,将其分为生成器和评论家角色,用辩论感知信号训练,通过LoRA微调少量参数。核心是反事实评论家优势,提高了模型准确率,在数学推理基准测试中有显著效果。
AI 中文摘要
大型语言模型虽能实现强大的推理性能,但训练成本高昂,这对预算有限下训练的紧凑模型(≤4B参数)来说挑战尤为严峻。我们引入了MADA-RL,这是一个训练后框架,将紧凑模型专门化为生成器和评论家角色,并使用辩论感知学习信号进行训练,仅通过LoRA适配器微调一小部分参数。我们的核心贡献是反事实评论家优势:一种动态的、基于角色的基线,将评论家的优势重新定义为其奖励减去生成器集合的实例准确率。这明确优化评论家以超越生成器共识,而非仅仅复制正确答案,比静态平均奖励归一化产生更有针对性的信用分配。在部署时,专门的智能体以轻量级多轮协议组合。在五个数学推理基准测试中,MADA-RL将DeepSeek-R1-Distill-Qwen-1.5B模型的准确率从39.9%提高到41.9%(提高2.0个百分点,p < 0.001),使用的可训练参数比完全微调的基线少16倍,使其处于准确率-可训练参数帕累托前沿。它接近但未超过最强基线(DeepScaleR,STILL-3),后者在大得多的数据集上训练;我们直接分析了这一差距和相关的推理时间成本。一项对照研究分离出MADA-RL收益的来源:反事实优势在所有评估模型中产生了最高的评论家改进率,表明训练后的评论家学会纠正生成器错误而非模仿它们。
英文摘要
Large language models achieve strong reasoning performance, but often at prohibitive training cost - a challenge that is especially acute for compact models ($\leq 4 \, \mathrm{B}$ parameters) trained under limited budgets. We introduce MADA-RL, a post-training framework that specializes compact models into generator and critic roles and trains them with a debate-aware learning signal, fine-tuning only a small subset of parameters via LoRA adapters. Our central contribution is a counterfactual critic advantage: a dynamic, role-conditioned baseline that redefines the critic's advantage as its reward minus the generator ensemble's per-instance accuracy. This explicitly optimizes critics to improve over generator consensus rather than to merely reproduce a correct answer, yielding more targeted credit assignment than static mean-reward normalization. At deployment, the specialized agents are composed in a lightweight multi-round protocol. Across five mathematical reasoning benchmarks, MADA-RL raises the accuracy of the DeepSeek-R1-Distill-Qwen-1.5B model from $39.9 \, \%$ to $41.9 \, \%$ ($+2.0$ points, $p < 0.001$) using $16$ times fewer trainable parameters than fully fine-tuned baselines, placing it on the accuracy-trainable-parameter Pareto front. It approaches, but does not surpass, the strongest baselines (DeepScaleR, STILL-3), which are trained on substantially larger datasets; we analyse this gap and the associated inference-time cost directly. A controlled study isolates the source of MADA-RL's gains: the counterfactual advantage produces the highest critic improvement rate of any model evaluated, indicating that trained critics learn to correct generator errors rather than to imitate them.
Comments20 pages, 3 figures, 9 tables, 2 algorithms, under review at TMLR