发表机构
Kenan-Flagler Business School; University of North Carolina at Chapel Hill(凯南-弗拉格勒商学院; 北卡罗来纳大学教堂山分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文批判性综述AI缩放研究中的资源分配证据,提出能力表面与资源包络框架,揭示评估指标、部署量等如何影响分配结论,并明确证据支持的比较需求。
AI 中文摘要
AI缩放研究日益评估将预训练模型与检索、搜索、验证、工具和交互相结合的系统。然而,在更大预算下获得更高分数本身并不能表明额外资源的最佳投入方向。本批判性综合综述探讨了报告的缩放结果何时支持资源分配决策。它比较了预训练、测试时计算、检索和智能体评估方面的证据,区分了被测程序的性能与在资源限制下可实现的最佳性能。综合结果表明,这些证据中反复出现三种不匹配:在选定答案之前就计算成功、部署系统不会拥有的信息以及比较中遗漏的成本。能力表面将性能表示为预算、机制和可用信息的函数。经过推导的分析示例展示了评估指标、部署量、选择规则和停止策略如何改变分配结论。资源包络为报告分数背后的任务、开发和运行时资源、信息访问及程序提供了结构化记录。将其应用于已发表的比较,说明了证据支持哪些结论以及哪些部署问题仍未解决。由此产生的框架规定了在可行系统之间进行选择所需的比较,并激励了关于分配规则跨任务和运行条件迁移的实验。它不提出通用缩放定律,也不从基准提升中推断通用智能。
英文摘要
AI scaling studies increasingly evaluate systems that combine a pretrained model with retrieval, search, verification, tools, and interaction. Yet a higher score under a larger budget does not by itself show where additional resources are best spent. This critical integrative review asks when a reported scaling result supports a resource-allocation decision. It compares evidence across pretraining, test-time computation, retrieval, and agent evaluation, distinguishing the performance of a tested procedure from the best performance achievable under a resource limit. The synthesis shows that three mismatches recur across this evidence: success counted before an answer is chosen, information a deployed system will not have, and costs left out of the comparison. A capability surface expresses performance as a function of budgets, mechanisms, and available information. Worked analytical examples show how the evaluation metric, deployment volume, selection rule, and stopping policy can alter an allocation conclusion. A resource envelope provides a structured record of the task, development and run-time resources, information access, and procedure behind a reported score. Its application to a published comparison illustrates which conclusions the evidence supports and which deployment questions remain unresolved. The resulting framework specifies the comparisons needed to choose among feasible systems and motivates experiments on the transfer of allocation rules across tasks and operating conditions. It does not propose a universal scaling law or infer general intelligence from benchmark gains.
Comments35 pages, 3 figures, 9 tables