A Training-Time Diagnostic for Generalization via the Log-Alignment Ratio
基于对数对齐比率的训练时泛化诊断
机构 * Essential AI
AI总结 提出对数对齐比率(LAR)作为参数-激活对齐的度量,通过捕捉训练中权重谱与激活谱的扩散来跟踪记忆与泛化的转换,并在grokking和语言模型预训练中预测泛化差距。
Comments 32 pages, 25 figures
作者
Large Language Models
基于对数对齐比率的训练时泛化诊断
机构 * Essential AI
AI总结 提出对数对齐比率(LAR)作为参数-激活对齐的度量,通过捕捉训练中权重谱与激活谱的扩散来跟踪记忆与泛化的转换,并在grokking和语言模型预训练中预测泛化差距。
Comments 32 pages, 25 figures
Comments include MegaMath-Web-Pro
Comments 15 pages, 5 figures
Comments ICLR 2022 + Updated Checkpoint Release
Comments Technical Report, 20 pages, 13 figures, 19 tables
Comments CVPR 2021 Oral
Comments TACL 2020; pre-MIT Press publication version; v5 has a random attention baseline
Comments ICCV 2019
Comments Accepted at ACL 2019 as long paper
Comments Improved skewing section and accompanying figures. Previous titles are "An Improved Relative Self-Attention Mechanism for Transformer with Application to Music Generation" and "Music Transformer"
Comments Appears in International Conference on Machine Learning, 2018. Code available at https://github.com/tensorflow/tensor2tensor
Comments ICML 2018
Comments NAACL 2018
Comments arXiv admin note: text overlap with arXiv:1706.03762
Comments accepted at EMNLP 2016, Workshop on Structured Prediction for NLP. Oral presentation