AI 中文总结
本研究通过定义轨迹级协议,在PyTorch与Zig编写的独立框架numbat上对Qwen3-0.6B进行LoRA临床微调,发现了17个单实现遗漏的软件故障,凸显跨栈验证的价值。
AI 中文摘要
神经网络训练存在一个“预言机问题”:一次运行可能正常收敛并产生可用模型,但其底层软件却计算了与规定不符的内容。几乎所有此类工作都在单一栈上运行,因此很少有独立的对象可供校验。本研究探讨独立实现的训练栈能否作为整个微调流水线的差分预言机,而非以往差分测试所针对的算子和推理路径。我们定义了一种轨迹级协议——包含共享规范、涵盖算术、模型加载、数据渲染和学习轨迹的交叉检查点,以及栈、编排和语言运行时的独立性分离——并将其应用于Qwen3-0.6B的LoRA适配,该适配基于168574个临床问答对,分别在PyTorch和numbat(一个用Zig编写的独立框架,通过其C接口由六种语言原生驱动)上完成。在覆盖完整epoch的42次配对评估中,两个栈的保留交叉熵平均差异为0.134%,且四次实现的epoch结束时差异在0.15%以内。此次比较发现了17个单实现开发遗漏的故障,其中2个值得关注的软件工程故障。对训练模型影响最大的故障位于数值核之外:临床文本渲染方式的不匹配使保留损失移动了0.15,约为其旁发现的算术故障的500倍。另有4个故障仅能从内存模型与前两种实现不同的语言中触发:调度器跨线程迁移工作、收集器对设备内存视而不见、所有权规则需要接口缺少的原语。实现多样性有多个维度,运行时是其中之一。
英文摘要
Neural network training has an oracle problem: a run can converge normally and yield a usable model while the software beneath it computes something other than specified. Almost all such work runs on one stack, so there is rarely anything independent to check against. We study whether independently implemented training stacks can serve as differential oracles for a whole fine-tuning pipeline, rather than the operators and inference paths that prior differential testing targets. We define a trajectory-level protocol -- a shared specification, cross-check points spanning arithmetic, model loading, data rendering and the learning trajectory, and a separation of independence of the stack, the orchestration and the language runtime -- and apply it to a LoRA adaptation of Qwen3-0.6B over 168,574 clinical question-answer pairs under PyTorch and under numbat, an independent framework written in Zig, driven natively and through its C interface from six languages. Across 42 paired evaluations spanning a full epoch the two stacks' held-out cross-entropy differs by 0.134% on average, and four implementations end the epoch within 0.15% of one another. The comparison exposed 17 faults that single-implementation development had missed, two of them notable for software engineering. The fault with the largest effect on the trained model lay outside the numerical kernels: a mismatch in how clinical text was rendered moved held-out loss 0.15, some 500 times more than the arithmetic faults found beside it. And four faults were reachable only from a language whose memory model differs from the first two implementations: a scheduler migrating work across threads, a collector blind to device memory, an ownership discipline needing a primitive the interface lacked. Implementation diversity has several axes, and the runtime is one.
Comments15 pages, 3 figures, 8 tables