Numbat:构建并验证一个自包含的机器学习技术栈
Numbat: Building and Verifying a Self-Contained Machine-Learning Stack
浏览论文内容
中文总结 AI 辅助
本文介绍Numbat,一个用Zig编写的自包含机器学习栈,通过多级差分验证确保正确性,并在COCO上训练YOLOv8m达到与参考实现相当的精度。
中文摘要 AI 辅助
机器学习系统几乎完全建立在少数几个大型的、由Python编排的框架之上,并继承了这些技术栈的工程成本:包含数百个版本耦合包的环境、用于部署的独立导出工具链,以及研究编写语言与产品交付语言之间的割裂。我们报告了numbat的构建与验证工作,这是一个用单一通用语言(Zig)编写、无第三方运行时依赖的机器学习技术栈。该技术栈涵盖张量计算、自动微分、神经网络模块、混合精度、多GPU训练、数据加载与监控;一个SDK通过超过1400个入口点的稳定、增量版本化的C ABI暴露其功能,并提供六种语言的绑定;其临床领域平面将监管要求编码为可执行的验收门,而非文档。验证这样一个技术栈是构建工作中更困难的一半:有缺陷的训练运行很少失败,而是悄然收敛到略差的模型。我们将一个广泛使用的参考实现视为可执行规范,并在五个层面与之进行验证,从算子梯度检查到针对同机参考运行的自动化轨迹门——这一安排在我们的配套研究中被形式化为轨迹级差分预言机。该协议揭示了十个静默配方分歧,我们对其机制和症状进行了编目。作为验收测试,我们从随机初始化开始,在COCO 2017上按完整的500轮计划训练了一个25.9M参数的YOLOv8m类检测器:导出权重在官方协议下得分为0.4956 mAP50-95,由参考栈自身的验证器评分(已发布端点为0.502),在相同硬件上单GPU步长时间持平。权重、每轮指标和完整运行清单均已发布。
英文摘要
Machine-learning systems are built almost exclusively on a few large Python-orchestrated frameworks, and they inherit those stacks' engineering costs: environments of hundreds of version-coupled packages, separate export toolchains for deployment, and the split between the language research is written in and the language products ship in. We report on the construction and verification of numbat, a machine-learning stack written in one general-purpose language (Zig) with no third-party runtime dependencies. The stack spans tensor computation, automatic differentiation, neural-network modules, mixed precision, multi-GPU training, data loading and monitoring; an SDK exposes it behind a stable, additively versioned C ABI of over 1,400 entry points, with bindings for six languages; and its clinical domain planes encode regulatory requirements as executable acceptance gates rather than documentation. Verifying such a stack is the harder half of building it: a defective training run rarely fails, it converges quietly to a slightly worse model. We treat a widely used reference implementation as an executable specification and verify against it at five levels, from operator gradient checks to an automated trajectory gate against a same-machine reference run - the arrangement our companion study formalizes as a trajectory-level differential oracle. The protocol surfaced ten silent recipe divergences, which we catalog with mechanisms and symptoms. As the acceptance test, we train a 25.9M-parameter detector of the YOLOv8m class from random initialization on COCO 2017 for the full 500-epoch schedule: the exported weights score 0.4956 mAP50-95 under the official protocol, scored by the reference stack's own validator (published endpoint 0.502), with single-GPU step time at parity on identical hardware. Weights, per-epoch metrics and the full run manifest are released.
发表机构
- CloudKites AI Lab(CloudKites人工智能实验室)
- Monash Business School, Monash University(莫纳什大学莫纳什商学院)
机构由 AI 辅助整理,请以论文原文为准。