发表机构
Adioris Tech Ltd.(Adioris科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本报告实测了用 Rust 端到端预训练语言模型的可行性,记录了 Candle 和 Burn 框架的静默缺陷及验证方法,并展示了以孟加拉语为主的模型效果,结论是 Rust 尚不适合训练但适合推理。
AI 中文摘要
我独自一人,在没有团队、没有 PyTorch、训练路径中也没有 Python 的情况下,用 Rust 端到端预训练了一个语言模型,租用 GPU 的费用为 164 美元。我将此作为一项成就来报告,而非一项推荐:更有用的贡献是对 2026 年两大领先 Rust 机器学习框架 Candle 和 Burn 作为训练(而非推理)后端时缺陷的实测分类。我记录了五个 Candle 缺陷,包括静默不产生梯度的融合内核,以及三个 Burn 缺陷,包括反向传播吞吐量仅为理论 GPU 吞吐量的约 3%,以及一个在数十亿参数规模下训练中途导致段错误的核融合路径。每一个缺陷都能通过常规损失曲线检查;没有一个会主动暴露自身。我描述了捕获这六个静默失败的验证纪律,其核心是一个梯度流仲裁器:一个运行一次前向/反向传播并断言每个可训练参数都收到有限非零梯度的测试,该测试可推广到任何框架。训练出的模型(约 0.4B 参数,以孟加拉语为主)显示出强烈的孟加拉语语言建模信号——每个 token 的负对数似然为 0.93,而随机初始化孪生模型为 12.60——但在英语常识多项选择上表现与随机猜测相当,这是刻意采用小规模、以孟加拉语为主的预算(约 20 亿 token,54.6 小时,租用一块 H100)的预期结果。我还报告了孟加拉文字中的分词器丰度陷阱:朴素的字节级分词将孟加拉语压缩到每 token 约 1.4 个字符,而英语为 3.9,静默地颠倒了语料库的语言平衡;修复后达到约 4.1。据我所知,这是首批有记录的纯 Rust 端到端 LM 预训练运行之一。在此次运行之后,我将训练迁移到 PyTorch,并保留 Rust 用于设备端推理:在我看来,Rust 尚不是训练语言模型的竞争性选择,尽管它可能是推理的好选择。
英文摘要
I pretrained a language model end-to-end in Rust - alone, with no team, no PyTorch, and no Python in the training path - for $164 in rented GPU time. I report that as an achievement, not a recommendation: the more useful contribution is a measured failure taxonomy of the two leading Rust ML frameworks, Candle and Burn, as training (not inference) backends in 2026. I document five Candle defects, including fused kernels that silently produce no gradient, and three Burn defects, including a backward pass at roughly 3% of theoretical GPU throughput and a kernel-fusion path that segfaults mid-training at multi-billion-parameter scale. Every one passed ordinary loss-curve inspection; none announced itself. I describe the verification discipline that caught six such silent failures, centered on a gradient-flow arbiter: a test that runs one forward/backward pass and asserts every trainable parameter receives a finite, nonzero gradient, generalizable to any framework. The trained model (roughly 0.4B parameters, Bangla-first) shows strong Bangla language-modeling signal - a per-token negative log-likelihood of 0.93 against 12.60 for a random-initialized twin - while scoring at chance on English commonsense multiple-choice, the expected outcome of a deliberately small, Bangla-weighted budget (about 2 billion tokens, 54.6 hours, one rented H100). I also report a tokenizer-fertility trap in Bengali script: naive byte-level tokenization collapsed Bangla to roughly 1.4 characters per token against English's 3.9, silently inverting the corpus's language balance; fixing it reached roughly 4.1. To my knowledge, this is among the first documented end-to-end LM pretraining runs in pure Rust. After this run I moved training to PyTorch and kept Rust for on-device serving: in my hands, Rust is not yet a competitive place to train a language model, though it may be a good place to serve one.