arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UMSS:迈向无监督多模态语义分割

UMSS: Towards Unsupervised Multi-modal Semantic Segmentation

Haitian Zhang, Thai Duy Nguyen, Xiangyuan Wang, Mohan Liu, Lin Wang

arXiv 2607.12372首次发表:更新:

发表机构

EmPACT Lab, School of EEE, Nanyang Technological University; The University of Hong Kong(电气与电子工程学院电磁脉冲与天线研究室,南洋理工大学; 香港大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对无监督多模态语义分割问题,提出基于DINOv3的UniM2框架,通过跨模态对应协同学习统一潜在空间提取语义线索,并引入跨模态协调器缓解冲突,实验证明该框架相比现有框架有明显优势。

AI 中文摘要

多模态语义分割对于复杂环境中的稳健感知至关重要,但由于人工标注成本高昂,其潜力尚未得到充分挖掘。无监督语义分割在单RGB模态上取得了不错成果,但其简单扩展到多模态数据常受融合退化阻碍。本文首次尝试解决无监督多模态语义分割这一全新问题,提出基于DINOv3的新型框架UniM2。通过由跨模态对应协同驱动学习统一潜在空间来提取内在共享语义线索,引入跨模态协调器以缓解模态冲突。在NYU Depth v2和MFNet上的实验结果表明,UniM2分别将平均交并比提高了6.4%和9.8%,优于现有框架。

英文摘要

Multimodal semantic segmentation (MSS) is essential for robust perception in complex environments, yet its potential remains largely untapped because of the prohibitive cost of human annotations. While unsupervised semantic segmentation (USS) has achieved strong results on a single RGB modality, its naive extension to multimodal data is often hindered by fusion degradation. This occurs because, without explicit supervision, existing frameworks struggle to reconcile the heterogeneous structural patterns captured by different sensors and therefore fail to effectively exploit their complementary information. In this paper, we make the first attempt to address the novel problem of Unsupervised Multimodal Semantic Segmentation (UMSS), aiming to effectively exploit complementary sensor information in a fully label free setting. To this end, we propose UniM2 (Unified Multimodal), a novel framework built on DINOv3 that transforms conventional fusion methods into consistent performance gains. Our key idea is to learn a unified latent space driven by Cross Modal Correspondence Synergy (CMCS) to extract intrinsic shared semantic cues, bypassing the need for label guided adaptive fusion. To mitigate inherent intermodal conflicts, we introduce a Cross Modal Harmonizer (CMH) that designates RGB as a stable reference, effectively suppressing inconsistent relational supervision while guiding the model to exploit complementary structural features. Extensive experimental results on NYU Depth v2 and MFNet show that UniM2 improves mIoU by 6.4% and 9.8%, respectively, demonstrating clear advantages over existing frameworks for UMSS.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑