大型推理模型在推理时间上的能力与效率缩放
Scaling of Capability and Efficiency at Inference Time in Large Reasoning Models
浏览论文内容
中文总结 AI 辅助
本研究通过层次贝叶斯模型分析DeepSeek-R1-Distill模型,发现推理能力随规模收益递减,而效率不随规模提升,揭示朴素缩放的局限性。
中文摘要 AI 辅助
能力与效率是大型语言模型(LLMs)推理中的两个关键维度。能力指正确解决给定问题的能力,而效率指在有限资源下完成该任务的能力。当LLMs使用思维链(CoT)推理来解决受控难度的问题时,正确解决的问题数量以及达到正确答案所需的令牌数量都取决于问题难度和模型规模。然而,这些因素如何共同塑造能力与效率仍知之甚少。在此,我们使用层次贝叶斯模型评估DeepSeek-R1-Distill模型家族中LLMs在四类算术和算法推理问题上的能力与效率。在固定模型规模下,正确解决实例的概率随实例大小(我们作为问题难度的代理)近似指数衰减。衰减尺度随模型规模次线性增长,表明更大的模型能力更强,但能力增益随规模递减。输出长度随实例大小(作为难度代理)呈幂律增长。然而,该幂律的参数并不随模型规模系统性变化,表明更大的模型并未变得更高效。总之,这些发现揭示了朴素缩放作为开发更强大AI系统策略的潜在局限性:能力提升收益递减,而效率几乎没有提升。
英文摘要
Capability and efficiency are two key dimensions of reasoning in large language models (LLMs). Capability refers to the ability to solve a given problem correctly, whereas efficiency refers to the ability to do so with limited resources. When LLMs use Chain-of-Thought (CoT) reasoning to solve problems of controlled hardness, both the number of problems solved correctly and the number of tokens required to reach a correct answer depend on problem hardness and model size. However, how these factors jointly shape capability and efficiency remains poorly understood. Here, we use hierarchical Bayesian models to evaluate the capability and efficiency of LLMs from the DeepSeek-R1-Distill model family across four classes of arithmetic and algorithmic reasoning problems. At a fixed model size, the probability of correctly solving an instance decays approximately exponentially with instance size, our proxy for problem hardness. The decay scale grows sublinearly with model size, indicating that larger models are more capable, but that capability gains diminish with scale. Output length grows as a power law with instance size, which serves as a proxy for difficulty. However, the parameters of this power law do not vary systematically with model size, suggesting that larger models do not become more efficient. Together, these findings reveal potential limitations of naive scaling as a strategy for developing more capable AI systems: capability improves with diminishing returns, while efficiency shows little to no improvement.
发表机构
- Northeastern University(东北大学)
- Complexity Science Hub Vienna(维也纳复杂性科学中心)
- Santa Fe Institute(圣塔菲研究所)
机构由 AI 辅助整理,请以论文原文为准。