发表机构
INM – Leibniz Institute for New Materials; Saarland University; German Research Center for Artificial Intelligence (DFKI); University of Stuttgart; NEC Laboratories Europe(莱布尼茨新材料研究所; 萨尔兰大学; 德国人工智能研究中心; 斯图加特大学; NEC欧洲实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FlashCart通过生成GPU内核和递归压缩架构,使高阶笛卡尔张量积计算高效,在SPICE-MACE-OFF上以560万参数超越1.89亿参数Transformer的精度和速度。
AI 中文摘要
机器学习原子间势将原子模拟扩展到电子结构方法可及的长度和时间尺度之外。然而,等变架构的计算成本限制了它们在实践中能够表示的局部相关性,从而限制了它们可达到的精度。在此,我们引入FlashCart,通过将生成的GPU内核与一种递归构建等变特征并在每一步将其压缩为固定宽度的架构相结合,使高阶相关性变得可负担。我们以独立的笛卡尔分量表达张量积,并对其及其导数进行符号简化,生成融合内核,这些内核通常优于优化的球谐对应物。然后我们表明,增加相关阶数比增加宽度、深度或张量秩更有效地提高精度。在SPICE-MACE-OFF上,FlashCart模型推进了实测的精度-效率前沿:一个拥有560万参数的模型实现了更低的能量和力误差,并且推理速度比拥有1.89亿参数的Transformer快10倍。
英文摘要
Machine-learned interatomic potentials extend atomistic simulations beyond the length- and timescales accessible to electronic-structure methods. However, the computational cost of equivariant architectures limits the local correlations they can represent in practice and therefore their achievable accuracy. Here we introduce FlashCart, which makes higher-order correlations affordable by combining generated GPU kernels with an architecture that recursively builds equivariant features and compresses them to a fixed width at each step. We express tensor products in independent Cartesian components and symbolically simplify them and their derivatives, producing fused kernels that often outperform optimized spherical counterparts. We then show that increasing correlation order improves accuracy more efficiently than increasing width, depth, or tensor rank. On SPICE-MACE-OFF, FlashCart models advance the measured accuracy-efficiency frontier: a model with $5.6$ million parameters achieves lower energy and force errors and $10\times$ faster inference than a transformer with $189$ million parameters.