AI 中文总结
针对时空耦合建模中灵活性与物理先验难以兼顾的问题,提出闵可夫斯基位置编码(MinkowskiPE),利用洛伦兹变换显式引入几何偏置,在分子动力学和视频预测任务上取得最优结果,且参数效率高。
AI 中文摘要
建模时空耦合是构建从微观到宏观尺度的物理智能的关键挑战。现有模型通过物理启发的动力学公式或学习驱动的架构来广泛捕获这种结构。前者提供更强的先验,但可能限制灵活性,而后者更灵活,但使时空耦合在很大程度上隐式化。因此,我们寻求一种将灵活学习与显式几何偏置相结合的方法,以联合建模时间和空间。为此,我们提出了闵可夫斯基位置编码(MinkowskiPE),它使用联合的时间和空间坐标来参数化应用于查询和键特征的洛伦兹变换。使用MinkowskiPE,查询-键注意力分数仅通过两个标记之间的相对时空位移依赖于位置,因此对坐标的全局平移不变。该范式保留了标准的点积注意力接口,并与高效的注意力实现兼容。我们在微观分子动力学和宏观视频预测任务上评估了MinkowskiPE,在所有九项多轨迹分子评估中取得了最佳结果,并将KTH视频预测MSE相对于最佳基线降低了9.9%,同时使用的参数约为其十分之一。
英文摘要
Modeling spatiotemporal coupling is a key challenge in building physical intelligence across scales, from microscopic to macroscopic. Existing models capture such structure broadly through physics-motivated dynamical formulations or learning-motivated architectures. The former provide stronger priors but may constrain flexibility, whereas the latter are more flexible but leave the spatiotemporal coupling largely implicit. We therefore seek an approach that combines flexible learning with an explicit geometric bias for jointly modeling time and space. To this end, we propose Minkowski Positional Encoding (MinkowskiPE), which uses joint temporal and spatial coordinates to parameterize Lorentz transformations applied to query and key features. With MinkowskiPE, the query-key attention score depends on position only through the relative spacetime displacement between the two tokens and is therefore invariant to global translation of the coordinates. This paradigm retains the standard dot-product attention interface and remains compatible with efficient attention implementations. We evaluate MinkowskiPE on microscopic molecular dynamics and macroscopic video prediction tasks, achieving the best results on all nine multi-trajectory molecular evaluations and reducing KTH video-prediction MSE by 9.9% relative to the best baseline while using roughly one-tenth as many parameters.
Comments21 pages, 4 figures, 4 tables