AI 中文总结
TriO利用相机、激光雷达和雷达三种模态进行无监督自监督,预测4D占用、障碍物分割、光流和激光雷达,实现开放集长尾物体分割,并在多个3D/4D任务上达到最先进水平。
AI 中文摘要
我们提出了TriO,一种多模态无监督世界模型,用于预测4D占用、障碍物分割、光流和激光雷达数据。与以往工作不同,TriO利用三种不同的传感器模态(相机、激光雷达和雷达)作为输入和自监督来源,无需额外的人工标注。得益于其新颖的监督方式,该模型能够从可行驶表面分割任何占用物,克服了现有开放集方法在处理长尾物体时的局限性。TriO在多个3D和4D任务中取得了最先进的结果,包括占用、光流和激光雷达预测,以及在Argoverse 2和Spotting the Unexpected等多个数据集上的零样本道路障碍物分割。
英文摘要
We present TriO, a multi-modal unsupervised world model that predicts 4D occupancy, obstacle segmentation, flow and LiDAR. In contrast to prior work, TriO utilizes three distinct sensor modalities (camera, LiDAR, and RADAR) as both inputs and sources of self-supervision, eliminating the need for additional human annotations. Thanks to its novel supervision, the model is able to segment any occupancy from the drivable surface, overcoming the limitations of existing open-set methods in handling long-tail objects. TriO achieves state-of-the-art results in multiple 3D and 4D tasks, including occupancy, flow, and LiDAR prediction, as well as zero-shot road obstacle segmentation across multiple datasets such as Argoverse 2, and Spotting the Unexpected.
CommentsPublished at ECCV 2026, 49 pages, 20 figures