发表机构
Stanford University; New York University(斯坦福大学; 纽约大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对机器学习验证难题,本文提出见证者方法,利用行为指纹与重放挑战审计训练,以极小开销实现可靠认证,并引入自动认证排行榜促进可复现性。
AI 中文摘要
机器学习的进步不能超越我们验证它的能力。随着如今论文数量的激增,每项科学主张最初都依赖于对训练者的信任,这导致了评估和基线的不均衡,并阻碍了可靠进展。传统上,验证的责任落在读者身上,读者必须复现昂贵的训练过程。由于低质量贡献的激增、方法的多样性以及所需计算量的庞大,这种策略变得不切实际。我们将举证责任放在其应属之处——训练者身上,并在此过程中显著降低了验证的总体成本。我们引入了见证者(Witnesses),一种用于认证神经网络训练运行中的数据使用、数据使用和评估的方法。我们的关键洞察是,快速的行为指纹配合偶尔的重放挑战足以审计神经网络训练。我们的方法可大规模应用,对训练者开销极小,对验证者成本低廉,能以可放大的概率拒绝不良训练运行,并允许对数据包含和排除进行精确查询。我们在从1亿到20亿参数规模的语言模型训练运行中,跨DDP和FSDP测试了我们的方法,并证明了这种最小开销。我们还引入了一个自我调节的“自动认证”训练运行排行榜,以实现共享基线和进展。我们邀请社区参与该排行榜,以提高机器学习中的可复现性。
英文摘要
Progress in machine learning cannot outpace our ability to verify it. With an explosion in papers today, every scientific claim rests initially on trust in the trainer, leading to uneven evaluation, baselines, and forestalling of reliable progress. Traditionally, the burden of verification falls on the reader, who must reproduce expensive training runs. This strategy is impractical due to an explosion in slop contributions, diversity of methods, and the sheer compute required. We put the burden of proof where it belongs, on the trainer, and in the process also cut the overall cost of verification significantly. We introduce Witnesses, a method for certifying training, data usage and evaluation in a neural network training run. Our key insight is that fast behavioral fingerprints with occasional replay challenges are sufficient for auditing neural network training. Our method is applicable at scale with minimal overhead to the trainer, is cheap for the verifier, rejects bad training runs with amplifiable probability, and allows for exact queries of both data inclusion and exclusion. We test our method on language model training runs from 100M to 2B scales, across DDP and FSDP, and demonstrate this minimal overhead. We also introduce a self-regulating leaderboard of "auto-certified" training runs that enables shared baselines and progress. We invite the community to participate in the leaderboard to improve reproducibility in machine learning.
Comments10 pages, 15 appendix pages, 17 figures