A Training-Time Diagnostic for Generalization via the Log-Alignment Ratio
基于对数对齐比率的训练时泛化诊断
机构 * Essential AI
AI总结 提出对数对齐比率(LAR)作为参数-激活对齐的度量,通过捕捉训练中权重谱与激活谱的扩散来跟踪记忆与泛化的转换,并在grokking和语言模型预训练中预测泛化差距。
Comments 32 pages, 25 figures
作者
Large Language Models
基于对数对齐比率的训练时泛化诊断
机构 * Essential AI
AI总结 提出对数对齐比率(LAR)作为参数-激活对齐的度量,通过捕捉训练中权重谱与激活谱的扩散来跟踪记忆与泛化的转换,并在grokking和语言模型预训练中预测泛化差距。
Comments 32 pages, 25 figures
Essential-Web v1.0:24万亿token的有组织网络数据
AI总结 本文提出Essential-Web v1.0,一个24万亿token的带十二类分类标注的网络数据集,通过轻量模型标注和SQL过滤,在数学、代码、STEM和医学任务上取得竞争力提升。
Comments include MegaMath-Web-Pro
Muon在预训练中的实际效率
AI总结 该研究证明二阶优化器Muon在预训练的计算-时间权衡上拓展了优于AdamW的帕累托前沿,结合muP实现高效超参数迁移,提出低开销伸缩算法,经40亿参数规模实验验证可实现更经济的大批量训练。
重新思考预训练中的反思
机构 * Essential AI
AI总结 该研究重新思考了预训练中的反思能力,通过在思维链中引入故意错误测试模型的自我纠正能力,发现这种能力在预训练早期即出现并稳步提升,如OLMo2-7B模型在六项自我反思任务上展现了自我纠正能力。
机构 * Google Brain(谷歌大脑) ; Google Research(谷歌研究院) ; University of Toronto(多伦多大学) ; Google(谷歌)
Comments 15 pages, 5 figures
机构 * Google Research(谷歌研究院)
机构 * Google Research & DeepMind(谷歌研究院与深度思维)
Comments ICLR 2022 + Updated Checkpoint Release
机构 * UC Berkeley(加州大学伯克利分校) ; Google Research(谷歌研究院)
Comments Technical Report, 20 pages, 13 figures, 19 tables
机构 * Google Research(谷歌研究院) ; UC Berkeley(加州大学伯克利分校)
Comments CVPR 2021 Oral
机构 * Language Technologies Institute, Carnegie Mellon University(卡内基梅隆大学语言技术研究所) ; Google Research(谷歌研究院)
机构 * Google Research(谷歌研究院)
Comments TACL 2020; pre-MIT Press publication version; v5 has a random attention baseline
机构 * Google Brain(谷歌大脑)
Comments ICCV 2019
机构 * Google Research(谷歌研究院)
Comments Accepted at ACL 2019 as long paper
机构 * Google Research(谷歌研究院)
机构 * Google Brain(谷歌大脑)
Comments Improved skewing section and accompanying figures. Previous titles are "An Improved Relative Self-Attention Mechanism for Transformer with Application to Music Generation" and "Music Transformer"
机构 * Google Brain(谷歌大脑)
机构 * DeepMind(深度思维) ; Google Brain(谷歌大脑) ; MIT(麻省理工学院) ; University of Edinburgh(爱丁堡大学)
机构 * Google Brain(谷歌大脑) ; University of California, Berkeley(加州大学伯克利分校) ; Google AI(谷歌AI)
Comments Appears in International Conference on Machine Learning, 2018. Code available at https://github.com/tensorflow/tensor2tensor
机构 * Google Brain(谷歌大脑)
Comments ICML 2018
机构 * Google(谷歌) ; Google Brain(谷歌大脑)
Comments NAACL 2018
机构 * Google Brain(谷歌大脑) ; DeepMind(深度思维)
Comments arXiv admin note: text overlap with arXiv:1706.03762
机构 * Google Brain(谷歌大脑) ; University of Toronto(多伦多大学) ; Google Research(谷歌研究院)
机构 * Information Sciences Institute, University of Southern California(南加州大学信息科学研究所) ; Informatics Institute, University of Amsterdam(阿姆斯特丹大学信息学研究所) ; Google Brain(谷歌大脑)
Comments accepted at EMNLP 2016, Workshop on Structured Prediction for NLP. Oral presentation