面向4D占用预测的几何感知时空上下文建模
Geometry-Aware Spatio-Temporal Context Modeling for 4D Occupancy Forecasting
浏览论文内容
中文总结 AI 辅助
针对4D占用预测现有方法的几何失真与时间一致性问题,提出GAST方法,经Occ3D-nuScenes实验较SOTA实现mIoU、IoU提升及加速,长期预测性能强。
中文摘要 AI 辅助
4D占用预测用于建模3D场景的时空演化,对自动驾驶至关重要,尤其适用于极端场景仿真。现有方法常依赖离散标记化后进行自回归预测,但存在静态结构几何失真、预测时域内时间一致性不一致的问题。本研究提出一种面向4D占用预测的几何感知时空上下文建模方法(GAST),基于渐进式显式-隐式生成与双路径时空建模构建。具体而言,生成模块通过姿态驱动的变形、运动感知特征调制及基于注意力的特征细化,生成具有高几何保真度和语义合理性的逐帧占用;随后,时空模块通过全局上下文聚合增强空间一致性,同时通过时间动态提取捕捉场景演化。该统一设计可端到端联合优化历史重建与未来预测。在Occ3D-nuScenes数据集上的大量实验表明,本方法优于现有SOTA,mIoU提升7.67%、IoU提升6.44%,同时实现2.84倍加速,且在长期预测中保持强性能。
英文摘要
4D occupancy forecasting models the spatio-temporal evolution of 3D scenes and is crucial for autonomous driving, especially for corner-case simulation. Existing methods often rely on discrete tokenization followed by autoregressive prediction, yet struggle with geometric distortion in static structures and inconsistent temporal coherence over the forecasting horizon. In this work, we propose a Geometry-Aware Spatio-Temporal context modeling method (GAST) for 4D occupancy forecasting, built upon progressive explicit-implicit generation and dual-path spatio-temporal modeling. Specifically, the generation module produces per-frame occupancy with high geometric fidelity and semantic plausibility through pose-driven warping, motion-aware feature modulation, and attention-based feature refinement. Subsequently, the spatio-temporal module enhances spatial consistency through global context aggregation while capturing scene evolution through temporal dynamics extraction. This unified design enables joint optimization of historical reconstruction and future forecasting in an end-to-end manner. Extensive experiments on Occ3D-nuScenes demonstrate the superiority of our method, outperforming the state-of-the-art by 7.67% in mIoU and 6.44% in IoU with a 2.84x speedup, while maintaining strong performance in long-term forecasting.
发表机构
- South China University of Technology(华南理工大学)
- Institute of Optics and Electronics, CAS(中国科学院光电技术研究所)
机构由 AI 辅助整理,请以论文原文为准。