4DGS-JEPA:动态高斯泼溅的时间组合联合嵌入预测
4DGS-JEPA: Temporally Compositional Joint-Embedding Prediction for Dynamic Gaussian Splatting
浏览论文内容
中文总结 AI 辅助
提出4DGS-JEPA,一种用于动态高斯场景因果多时域预测的联合嵌入架构,通过时间组合原则和混合对应机制,实现无需完整未来重建的可复用预测动力学。
中文摘要 AI 辅助
动态高斯泼溅提供了演化三维场景的显式表示,但现有方法主要针对重建、未来状态生成或渲染进行优化,而非用于学习可复用的预测性动力学。我们提出4DGS-JEPA,一种高斯原生的联合嵌入预测架构,用于动态高斯场景上的因果多时域预测。该模型使用分层场景级、运动组级和高斯级表示,并结合一个时域条件转换算子,支持直接预测和递归展开。其核心原则是时间组合:到达同一未来端点的不同时间转换路径应产生兼容的预测状态。端点和多时域路径监督将这些预测锚定到未来目标嵌入,而选择性几何解码器和几何级组合将学习到的动力学锚定在一致的运动组运动和高斯几何中,无需完整的未来外观重建。我们进一步引入一种混合对应机制,该机制在重排序和拓扑变化下将持久规范身份与残差最优传输匹配相结合。我们在理论上刻画了零损失路径一致性和有限误差展开累积。三个受控实验提供了机制层面的证据,表明时间组合减少了潜在路径依赖性,同时保持了预测准确性;几何级组合提高了解码运动的一致性;混合对应在对应关系变得模糊时保持可靠身份,同时保持鲁棒性。总之,4DGS-JEPA为动态高斯世界提供了一种预测性、时间组合的表述。
英文摘要
Dynamic Gaussian Splatting provides an explicit representation of evolving 3D scenes, but existing approaches are primarily optimized for reconstruction, future-state generation, or rendering rather than for learning reusable predictive dynamics. We propose 4DGS-JEPA, a Gaussian-native joint-embedding predictive architecture for causal multi-horizon prediction over dynamic Gaussian scenes. The model uses a hierarchical scene-, motion-group-, and Gaussian-level representation together with a horizon-conditioned transition operator that supports both direct prediction and recursive rollout. Its central principle is temporal composition: different chronological transition paths reaching the same future endpoint should produce compatible predictive states. Endpoint and multi-horizon path supervision anchor these predictions to future target embeddings, while a selective geometry decoder and geometry-level composition ground the learned dynamics in consistent group motion and Gaussian geometry without requiring complete future appearance reconstruction. We further introduce a hybrid correspondence mechanism that combines persistent canonical identity with residual optimal-transport matching under reordering and topology change. We characterize zero-loss path agreement and finite-error rollout accumulation theoretically. Three controlled experiments provide mechanism-level evidence that temporal composition reduces latent path dependence while retaining predictive accuracy, geometry-level composition improves consistency of decoded motion, and hybrid correspondence preserves reliable identity while remaining robust when correspondence becomes ambiguous. Together, 4DGS-JEPA provides a predictive, temporally compositional formulation of dynamic Gaussian worlds.