arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

为精度而训练,为规模而执行:等变原子基础模型的架构保持推理

Train for Accuracy, Execute at Scale: Architecture-Preserving Inference for Equivariant Atomistic Foundation Models

Lei Fu, Zihui Feng, Yongheng Li, Hongwei Du, Xin He, Junyi Wu, Kejie Bao, Yueyu Zhang, Zeyu Deng, Ziheng Lu, Bonan Zhu

arXiv 2610.01036首次发表:更新:

发表机构

Beijing Institute of Technology; Zhongguancun Academy; Zhongguancun Institute of Artificial Intelligence; Kairos Materials; National University of Singapore(北京理工大学; 中关村学院; 中关村人工智能研究院; Kairos Materials; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Symmetrix-XL推理引擎,在不修改权重的情况下扩展MACE等变原子基础模型,通过流式边执行和专用代码生成,在A100上实现3.1-5.0倍加速并将容量提升至1124万原子,确立训练后执行作为新的扩展维度。

AI 中文摘要

等变原子基础模型提供了基于量子力学参考数据训练的广泛可迁移的原子间势,但其在模拟规模下的重复执行仍然计算和内存密集。我们提出了Symmetrix-XL,一个推理引擎,无需重新训练、蒸馏或修改学习权重即可扩展预训练的MACE检查点。它结合了流式边执行以避免图范围内扩展边中间体的物化,模型专用代码生成以编译检查点特定算子,以及分块执行以复用有界设备内存工作区。在完整的LAMMPS步基准上,Symmetrix-XL相对于ML-IAP + cuEquivariance在测试的A100和RTX 5090工作负载上将推理时间减少了3.1-5.0倍。在单个A100 80 GB GPU上,测试的MACE-OMAT-0容量边界从24,565个原子增加到1124万个原子。这大幅降低了模拟的硬件门槛,否则这些模拟需要跨多个GPU和计算节点进行空间分解。同一后端在64个A800 GPU上以93.8%的效率弱扩展到7.03亿个原子。能量、力、应力、分子动力学稳定性和Matbench Discovery评估在测量容差内再现了参考MACE行为。涵盖固态、界面和反应系统的案例研究展示了在实际工作流程中增加的能力。这些结果确立了训练后执行作为与模型重新设计、压缩和分布式扩展互补的扩展轴,并表明实际精度-成本前沿取决于模型架构和推理执行两者。

英文摘要

Equivariant atomistic foundation models provide broadly transferable interatomic potentials trained against quantum-mechanical reference data, but their repeated execution at simulation scale remains computationally and memory intensive. We present Symmetrix-XL, an inference engine that scales pretrained MACE checkpoints without retraining, distillation, or modification of their learned weights. It combines streamed-edge execution to avoid graph-wide materialization of expanded edge intermediates, model-specialized code generation to compile checkpoint-specific operators, and tiled execution to reuse a bounded device-memory workspace. On complete LAMMPS-step benchmarks, Symmetrix-XL reduces inference time by 3.1-5.0 times relative to ML-IAP + cuEquivariance across tested A100 and RTX 5090 workloads. On a single A100 80 GB GPU, the tested MACE-OMAT-0 capacity boundary increases from 24,565 atoms to 11.24 million atoms. This substantially lowers the hardware threshold for simulations that would otherwise require spatial decomposition across many GPUs and compute nodes. The same backend weak-scales to 703 million atoms on 64 A800 GPUs at 93.8 percent efficiency. Energy, force, stress, molecular-dynamics stability, and Matbench Discovery evaluations reproduce reference MACE behavior within measured tolerances. Case studies spanning solid-state, interfacial, and reactive systems demonstrate the increased capacity in realistic workflows. These results establish post-training execution as a scaling axis complementary to model redesign, compression, and distributed scale-out, and show that the practical accuracy-cost frontier depends on both model architecture and inference execution.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑