arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09444cs.CV

LighTROcc:基于实例中心三维高斯的轻量级4D占用预测

LighTROcc: Lightweight 4D Occupancy Forecasting via Instance-Centric 3D Gaussians

Hwanhee Jung, SeungHyeon Kim, Inkyu Koo, Qixing Huang, Sang Ho Yoon, Sangpil Kim

首次发表
浏览论文内容

中文总结 AI 辅助

LighTROcc提出轻量级实例中心框架,用紧凑查询和三维高斯表示可移动对象,单次前向预测当前与未来占用,在nuScenes上实现高效且时间一致的4D占用预测。

中文摘要 AI 辅助

从环视摄像头预测未来的三维占用对于自动驾驶至关重要,然而现有方法依赖于密集体素或鸟瞰图表示,其计算成本随空间分辨率和预测时域迅速增长。由于这些表示不显式维护对象身份,它们也难以随时间保持实例一致性。我们提出LighTROcc,一个轻量级的实例中心框架,用一组紧凑的可学习查询表示可移动对象,并在单次前向传播中预测当前和未来的占用。LighTROcc通过注意力引导的前向提升定位每个查询,结合图像空间交叉注意力、查询特定深度和相机几何来估计其三维中心。每个实例被建模为各向异性三维高斯的混合,并利用预测的位移在未来时间步中传播,产生连续且时间一致的占用预测。在nuScenes和补充的nuScenes-Occupancy上的实验表明,LighTROcc在实例级预测准确性上优于评估的密集和实例级基线,同时保持强体素级占用质量。在不同的模型配置中,LighTROcc在预测准确性和计算效率之间取得了有利的平衡,展示了紧凑实例中心建模用于基于相机的4D占用预测的潜力。

英文摘要

Forecasting future 3D occupancy from surround-view cameras is essential for autonomous driving, yet existing approaches rely on dense voxel or bird's-eye-view representations whose cost grows rapidly with spatial resolution and prediction horizon. Because these representations do not explicitly maintain object identities, they also struggle to preserve instance consistency over time. We present LighTROcc, a lightweight instance-centric framework that represents movable objects with a compact set of learned queries and predicts present and future occupancy in a single forward pass. LighTROcc localizes each query through attention-guided forward lifting, combining image-space cross-attention, query-specific depth, and camera geometry to estimate its 3D center. Each instance is modeled as a mixture of anisotropic 3D Gaussians and propagated across future steps using predicted displacements, producing continuous, temporally consistent occupancy forecasts. Experiments on nuScenes and supplemented nuScenes-Occupancy show that LighTROcc outperforms the evaluated dense and instance-wise baselines in instance-level forecasting accuracy while maintaining strong voxel-level occupancy quality. Across different model configurations, LighTROcc achieves a favorable balance between forecasting accuracy and computational efficiency, demonstrating the potential of compact instance-centric modeling for camera-based 4D occupancy forecasting.

发表机构

  • Korea University(高丽大学)
  • University of Texas at Austin(德克萨斯大学奥斯汀分校)
  • Korea Advanced Institute of Science & Technology(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

↑