用于垃圾分类语义分割无监督域适应的视觉语言引导伪标签
Vision-Language-Guided Pseudo-Labels for Unsupervised Domain Adaptation in Semantic Segmentation for Waste Sorting
浏览论文内容
中文总结 AI 辅助
该研究针对垃圾分类语义分割无监督域适应的标注难题,提出基于SAM、EVA-CLIP及可选BLIP的跨模态伪标签流水线,在两类域偏移场景下提升了性能,证实伪标签质量是自训练的关键。
中文摘要 AI 辅助
在自动驾驶、工业垃圾分类等应用场景中,获取语义分割的标注数据成本高昂且难以大规模实现。本文提出一种跨模态伪标签流水线,可在无任何目标域标注的情况下实现无监督域适应。该流水线基于两大核心基础模型构建:SAM生成与类别无关的区域提议,EVA-CLIP基于区域-文本相似度分配语义标签,置信度过滤确保仅可靠的伪标签用于自训练分割模型。作为可选扩展,BLIP为模糊区域提供语言基础的验证,无需改变整体流水线即可提升伪标签质量。在合成到真实自动驾驶、重点关注的实验室到工厂工业垃圾分类这两种域偏移场景下评估,该流水线始终优于仅使用源域的基线方法。结果表明,域偏移下自训练的关键因素是伪标签质量而非数量,跨模态语言基础为部署关键应用的可靠自动标注提供了可行路径。
英文摘要
Obtaining labeled data for semantic segmentation in applied settings (e.g., autonomous driving, industrial waste sorting) is expensive and often infeasible at scale. We present a cross-modal pseudo-labeling pipeline that enables unsupervised domain adaptation without any target-domain annotations. The pipeline is built on two core foundation models: SAM generates class-agnostic region proposals, and EVA-CLIP assigns semantic labels based on region-text similarity, with confidence filtering ensuring that only reliable pseudo-labels are used for self-training a segmentation model. As an optional extension, BLIP provides language-grounded verification for ambiguous regions, thereby improving pseudo-label quality without altering the overall pipeline. Evaluated on two domain shifts, synthetic-to-real autonomous driving and, with a primary focus, lab-to-factory industrial waste sorting, the pipeline consistently improves over source-only baselines. Our results demonstrate that pseudo-label quality, not quantity, is a decisive factor in self-training under domain shift, and that cross-modal language grounding offers a practical path to reliable automatic annotation in deployment-critical applications.
发表机构
- LMU Munich(慕尼黑大学)
- Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
- Fraunhofer Institute for Integrated Circuits IIS(弗劳恩霍夫集成电路研究所)
机构由 AI 辅助整理,请以论文原文为准。