arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OPEN-1B:一次完全可审计的训练运行

OPEN-1B: A Fully Auditable Training Run

John Donaghy, Brian Wilcox, Oğuzhan Ersoy, Shikhar Rastogi, Adam St Arnaud, Alexey Titov, Jordan Greenberg, Ben Fielding, Harry Grieve

arXiv 2609.17380首次发表:更新:

发表机构

Gensyn(Gensyn)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对开源语言模型无法证明可复现的问题,提出完全可审计的训练机制,通过确定训练非确定性顺序实现异构硬件上的比特级复现,并发布Open-1B模型及全套审计工具。

AI 中文摘要

开源语言模型存在可复现性问题。尽管发布了权重、训练数据和配方,但由于浮点运算的非结合性,它们都无法被证明是可复现的。深度学习框架通常提供确定性执行模式,允许在同一台机器上进行可复现的操作。不幸的是,这种确定性无法跨硬件传递,因此用户无法验证发布的检查点是否确实使用声明的训练配方生成。这为未披露的数据、注入的偏见或后门留下了空间,而现有的技术如学习证明或训练数据证明无法排除这些可能性。我们引入了模型透明度的新层级——完全可审计,在此层级下,训练期间对每个数据样本的每次操作都可以在异构商用硬件上以比特级确定性独立复现。通过对训练非确定性的来源(GPU内核归约、数据并行集群中的批次排序以及节点内/节点间集合通信)施加明确顺序,我们使得在单个商用硬件上重放大型分布式训练运行的任何单个步骤,并对照已发布的轨迹进行检查成为可能。由于在一台机器上重放整个运行不可行,我们通过一种集体验证方案来支持这一点,在该方案中,许多独立的审计员各自认证单个步骤,共同覆盖整个运行。我们发布了Open-1B,一个在此机制下训练的模型,连同其完整的预训练数据集、每个中间检查点、训练代码库以及复现和验证其训练任何步骤所需的审计工具。

英文摘要

Open-source language models have a reproducibility problem. Despite releasing weights, training data, and recipes, none of them are provably reproducible due to the non-associativity of floating-point arithmetic. Deep learning frameworks often offer a deterministic execution mode, allowing reproducible operations on the same machines. Unfortunately, this determinism does not carry across hardware such that a user can verify that a released checkpoint was actually produced using the declared training recipe. This leaves room for undisclosed data, injected biases, or backdoors that existing techniques such as proof-of-learning or proof-of-training-data cannot rule out. We introduce a new tier of model transparency, fully auditable, in which every operation on every data sample during training is independently reproducible on heterogeneous commodity hardware with bitwise certainty. By imposing a definite order on the sources of training nondeterminism, GPU kernel reductions, data batch ordering across a data-parallel cluster, and inter/intra-node collective communication, we make it possible to replay any individual step of a large, distributed training run on a single piece of commodity hardware and check it against the published trajectory. Because replaying an entire run on one machine is infeasible, we support this with a collective verification scheme in which many independent auditors each certify individual steps, together covering the whole run. We release Open-1B, a model trained under this regime, together with its full pretraining dataset, every intermediate checkpoint, the training codebase, and the audit harness needed to reproduce and verify any step of its training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑