AI 中文总结
该研究提出Meganeura编译器,通过Vulkan和Metal实现消费级GPU上的便携式训练与推理,在多数场景下性能接近或优于PyTorch,且编译速度更快、二进制文件更小,可支持跨平台部署。
AI 中文摘要
训练和已部署的推理通常需要跨越导出、转换以及特定平台的运行时边界。Meganeura探究一个紧凑的原生编译器能否在消费级GPU上同时覆盖训练与推理两个阶段。它具备类型化静态图、自动微分、优化器、检查点、内存规划器和运行时组件,可通过Vulkan和Metal将专用程序进行底层实现。我们在NVIDIA和AMD独立GPU、AMD APU、Apple硅以及Intel iGPU上,将五个匹配的工作负载与PyTorch进行对比。该协议将严格的f32精度与经过验证的快速路径分离,并分别控制前向和反向传播。50个设备-工作负载-模式单元中,有48个通过了两项控制;其余两个在新支持的APU上存在一处未解决的反向引用分歧。在严格的f32精度下,Meganeura在20个GPU参考的最小延迟单元中胜出12个,且有效训练的中位数差距为1.8倍。在AMD独立GPU上,五个推理工作负载中有四个与编译后的ROCm PyTorch性能差距在1.10倍以内,三个训练工作负载速度更快。在加速合约下,最大训练差距为4.6倍。编译耗时为0.1-2.4秒,而支持的GPU路径下的this http URL则需6-96秒; stripped二进制文件大小为13 MiB。调度分析将最大差距定位在卷积导数和注意力反向传播上。一项Android XR实际案例研究将Meganeura训练的解码器转移到共享图形队列的Adreno/OpenXR应用中。结果表明,通用消费级图形API可支持紧凑的训练到部署共享栈,且性能达到有用水平,有时可与厂商产品竞争。测得的差距指向内核覆盖、调度和算术策略,而非已识别的API限制。
英文摘要
Training and deployed inference often cross export, conversion, and platform-specific runtime boundaries. Meganeura asks whether one compact native compiler can span both phases on consumer GPUs. Its typed static graph, automatic differentiation, optimizer, checkpoint, memory planner, and runtime lower specialized programs through Vulkan and Metal. We compare five matched workloads with PyTorch on NVIDIA and AMD discrete GPUs, an AMD APU, Apple silicon, and an Intel iGPU. The protocol separates strict f32 from validated fast paths and gates forward and backward independently. Forty-eight of 50 device-workload-mode cells pass both gates; the other two share one unresolved backward-reference disagreement on a newly supported APU. In strict f32, Meganeura wins 12 of 20 GPU-referenced minimal-latency cells and has a median valid training gap of 1.8x. On the discrete AMD GPU, four of five inference workloads are within 1.10x of compiled ROCm PyTorch and three training workloads are faster. Under accelerated contracts, the worst training gap is 4.6x. Compilation takes 0.1-2.4 seconds versus 6-96 seconds for torch.compile on supported GPU paths; the stripped binary is 13 MiB. Dispatch profiles localize the largest gaps to convolution derivatives and attention backward. A physical Android XR case study transfers a Meganeura-trained decoder into an Adreno/OpenXR application sharing the graphics queue. The results show that general consumer graphics APIs can support a compact shared train-to-deploy stack at useful, sometimes vendor-competitive performance. The measured gaps point to kernel coverage, scheduling, and arithmetic policy rather than an identified API limitation.
Comments18 pages, 4 figures, 10 tables