发表机构
University of Maryland, College Park; Capital One(马里兰大学帕克分校; 第一资本)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出思维阶梯框架,通过渐进式问题改写与自我进化课程,显著提升中小规模语言模型的数学与多跳推理能力,优于知识蒸馏。
AI 中文摘要
大型语言模型(LLMs)在扩展到数千亿参数时擅长推理,但中小规模模型即使经过知识蒸馏(KD)仍然是脆弱的推理者。我们提出了思维阶梯(Ladders-of-Thought,LoT),一个通过结合渐进式问题改写与自我进化课程来提升推理能力的框架。LoT自动生成语义忠实但更易解决的推理问题变体,使用基于步骤的度量将其组织成难度桶,并采用自我进化的老虎机调度器自适应地分配训练。在两个推理领域(数学和多跳推理)上,跨不同家族的1-8B模型进行评估,LoT持续优于KD。它在算术任务上带来巨大提升(例如,在AddSub上+32个百分点,在SVAMP上+25个百分点),在域内测试分割上提升+2-8个百分点,并在多跳推理上表现出强但依赖数据集的效果(例如,在QASC上+16个百分点,在StrategyQA上+25个百分点)。LoT也比分阶段课程收敛更快,突显了自适应渐进的价值。这些结果表明,渐进式改写与自适应课程相结合,为增强较小LLM的推理能力提供了一种简单而有效的配方。
英文摘要
Large language models (LLMs) excel at reasoning when scaled to hundreds of billions of parameters, but small- and mid-scale models remain brittle reasoners even with knowledge distillation (KD). We present Ladders-of-Thought (LoT), a framework that improves reasoning by combining progressive question rewrites with a self-evolving curriculum. LoT automatically generates semantically faithful but easier variants of reasoning problems, organizes them into difficulty buckets using step-based measures, and employs a self-evolving bandit scheduler to allocate training adaptively. Evaluated on two reasoning domains, math and multi-hop reasoning, across 1-8B models from different families, LoT consistently improves over KD. It delivers large gains on arithmetic tasks (e.g., +32 percentage points on AddSub, +25pp on SVAMP), +2-8pp improvements on in-domain test splits, and strong though dataset-dependent benefits on multi-hop reasoning (e.g., +16pp on QASC, +25pp on StrategyQA). LoT also converges faster than staged curricula, highlighting the value of adaptive progression. These results show that progressive rewrites coupled with adaptive curricula provide a simple yet effective recipe for strengthening reasoning in smaller LLMs.