arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27588cs.LGcs.AI

能力流形与机器学习缩放定律

The Capability Manifold and ML Scaling Laws

  • University of Leeds(利兹大学)

机构由 AI 辅助整理,请以论文原文为准。

Syed Ali Raza Zaidi, Maryam Hafeez

AI总结:

本文提出能力流形框架,将下游能力映射到预训练、后训练和测试时资源,统一现有缩放定律,并量化能力对资源的敏感性。

AI中文摘要:

现有的机器学习(ML)缩放定律将预测损失与计算量、模型参数和数据相关联。然而,随着模型越来越多地通过智能体框架进行部署,仅凭损失不足以表征下游性能:具有相似损失的模型在推理、规划、检索和适应方面可能表现出不同的能力。然而,目前尚无统一框架将这些能力与机器学习生命周期中可用的耦合资源联系起来。我们通过引入能力流形来弥合这一差距,这是一个多维框架,通过有界缩放函数将下游能力映射到预训练、后训练和测试时资源。解析雅可比矩阵量化了能力对资源变化及其相互作用的敏感性。作为初步应用,我们将Kaplan型和Chinchilla型缩放定律以及测试时计算嵌入该框架,展示了现有的缩放关系如何统一为公共能力流形上的轨迹。

英文摘要:

Existing machine learning (ML) scaling laws relate predictive loss to compute, model parameters, and data. However, as models are increasingly deployed through agentic harnesses, loss alone is insufficient to characterize downstream performance: models with similar loss can exhibit different capabilities in reasoning, retrieval, planning, and adaptation. Yet, no unified framework connects such capabilities to the coupled resources available across the ML lifecycle. We bridge this gap by introducing a capability manifold, a multidimensional framework mapping downstream capabilities to pre-training, post-training, and test-time resources through bounded scaling functions. Analytical Jacobians quantify capability sensitivity to resource changes and interactions. As an initial application, we embed Kaplan- and Chinchilla-type scaling laws and test-time compute within the framework, demonstrating how existing scaling relationships can be unified as trajectories on a common capability manifold.

↑