发表机构
CUHKSZ; Tsinghua University; Shenzhen Polytechnic University(香港中文大学(深圳); 清华大学; 深圳职业技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出弱监督三维占用框架LetOccVote,利用跨帧投票改进几何与语义监督,在Occ3D-nuScenes数据集上实现二维伪标签监督方法中的最优性能。
AI 中文摘要
弱监督三维占用预测通过从视觉基础模型生成的二维伪标签中学习,减少了对成本高昂的三维标注的依赖。然而,现有方法通常直接将这些不完善的伪标签用作监督,使得占用学习容易受到错误的几何和语义目标的影响。我们观察到,重复观测之间的一致性为评估伪标签的可靠性提供了一种廉价且可靠的线索。基于这一观察,我们提出了LetOccVote,这是一种基于高斯的弱监督占用框架,利用跨帧投票来改进几何和语义监督。对于几何,Depth Vote利用跨帧几何一致性来细化支持的伪深度,并在体积提升和深度监督之前拒绝矛盾的估计;对于语义,Semantic Vote在共享三维空间中聚合伪语义观测,以识别可靠和有争议的证据,在加强可靠语义监督的同时过滤不可靠的伪标签片段。整个框架仅使用二维伪标签监督进行训练,不需要三维占用标注。在Occ3D-nuScenes数据集上,LetOccVote实现了53.27 IoU和20.39 mIoU,在使用二维伪标签监督的方法中达到了最先进的性能。
英文摘要
Weakly supervised 3D occupancy prediction reduces the reliance on costly 3D annotations by learning from 2D pseudo-labels generated by vision foundation models. However, existing methods typically use these imperfect pseudo-labels directly as supervision, making occupancy learning vulnerable to erroneous geometric and semantic targets. We observe that agreement across repeated observations provides an inexpensive and reliable cue for assessing pseudo-label reliability. Based on this observation, we propose \textbf{LetOccVote}, a weakly supervised Gaussian-based occupancy framework that leverages cross-frame voting to improve both geometric and semantic supervision. For geometry, Depth Vote exploits cross-frame geometric agreement to refine supported pseudo depth and reject contradictory estimates before volumetric lifting and depth supervision. For semantics, Semantic Vote aggregates pseudo-semantic observations in a shared 3D space to identify reliable and contested evidence, strengthening reliable semantic supervision while filtering unreliable pseudo-label segments. The entire framework is trained solely with 2D pseudo-label supervision without requiring 3D occupancy annotations. On Occ3D-nuScenes, LetOccVote achieves 53.27 IoU and 20.39 mIoU, establishing state-of-the-art performance among methods with 2D pseudo-label supervision.