arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ScAn-Bench:评估缩放分析方法论

ScAn-Bench: Evaluating Scaling Analysis Methodology

Artin Sermaxhaj, Nastaran Alipour, Donat Sinani, Johannes Hog, Neeratyoy Mallik, Jenia Jitsev, Danny Stoll

arXiv 2609.35707首次发表:更新:

发表机构

University of Freiburg; Zuse School ELIZA; Juelich Supercomputing Center (JSC), Research Center Juelich (FZJ)(弗莱堡大学; 楚泽学院ELIZA; 于利希超算中心(JSC),于利希研究中心(FZJ))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对缩放分析方法缺乏系统评估的问题,提出基于大量检查点的替代基准,首次系统评估跨模态的数据获取与外推方法论。

AI 中文摘要

机器学习的最新进展由大规模基础模型驱动,其中缩放定律以及为架构、数据和超参数寻找最优缩放方案是推进最先进水平的关键。因此,令人惊讶的是,尚无系统性研究评估在不同模型类型中获得缩放定律和方案的方法论。为揭示这一关键盲点并促进未来研究,我们基于语言和视觉-语言模型流程的4524个和8024个检查点,引入了替代基准ScAn-Bench-LLM和ScAn-Bench-VLM。在我们的基准上,我们首次对不同数据模态的缩放分析的数据获取和外推方法论进行了系统性评估。

英文摘要

Recent progress in machine learning is driven by large-scale foundation models, where scaling laws and finding optimal scaling prescriptions for architecture, data, and hyperparameters are key in advancing the state-of-the-art. Therefore, it is surprising that no systematic study evaluates the methodology to obtain scaling laws and prescriptions across different model types. To shed light on this crucial blind spot and facilitate future research, we introduce the surrogate benchmarks ScAn-Bench-LLM and ScAn-Bench-VLM based on 4524 and 8024 checkpoints of language and vision-language model pipelines. On our benchmarks, we perform the first systematic evaluation of both data acquisition and extrapolation methodology for scaling analysis across different data modalities.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑