感知锚定:用于无训练开放词汇语义分割的原型引导文本校准
Perceptual Anchoring: Prototype-Guided Text Calibration for Training-free Open-Vocabulary Semantic Segmentation
AI总结:
该研究针对无训练开放词汇语义分割的语义差距问题,提出原型引导文本校准(PTC),通过构建视觉原型校准文本嵌入,提升了多种现有方法的分割性能。
AI中文摘要:
无训练开放词汇语义分割(OVSS)可基于任意文本描述将图像划分为语义不同的区域,无需学习任何额外参数。然而,现有方法通常专注于改进视觉表征,而将仅编码通用类别概念的文本嵌入视为固定分类参考。这些通用概念与捕捉目标实例特定外观的视觉表征之间存在的语义差距,常导致不完整的掩码和非目标区域的错误预测。受感知锚定背后的符号-感知对应关系启发,我们提出用于无训练OVSS的原型引导文本校准(PTC)。在感知阶段,PTC基于初始匹配分数选择可靠的视觉证据,以构建特定类别的视觉原型;在锚定阶段,PTC利用这些原型校准对应的文本嵌入,校准强度根据视觉证据的数量自适应调整。因此,校准后的文本嵌入能更准确地与实例特定的视觉表征对齐,同时保留通用类别语义和开放词汇泛化能力。此外,PTC既无需额外训练也无需外部模型,可作为即插即用模块应用于现有方法。在八个基准上进行的大量实验表明,PTC显著提升了六种代表性方法的性能,并生成更完整、准确的分割结果。这些结果验证了PTC是一种改进视觉-文本对齐的简单有效方法。
英文摘要:
Training-free open-vocabulary semantic segmentation (OVSS) partitions an image into semantically distinct regions based on arbitrary text descriptions, without learning any additional parameters. However, existing methods typically focus on improving visual representations while treating text embeddings that encode only generic category concepts as fixed classification references. The resulting semantic gap between these generic concepts and the visual representations that capture the specific appearances of target instances often causes incomplete masks and erroneous predictions in non-target regions. Inspired by the symbol-percept correspondence underlying perceptual anchoring, we propose Prototype-Guided Text Calibration (PTC) for training-free OVSS. In the Perceiving stage, PTC selects reliable visual evidence based on initial matching scores to construct category-specific visual prototypes. In the Anchoring stage, PTC uses these prototypes to calibrate their corresponding text embeddings, with the calibration strength adaptively adjusted based on the amount of visual evidence. Consequently, the calibrated text embeddings align more accurately with instance-specific visual representations while preserving generic category semantics and open-vocabulary generalization. Moreover, PTC requires neither additional training nor external models and can serve as a plug-and-play module for existing methods. Extensive experiments across eight benchmarks show that PTC significantly enhances the performance of six representative methods and yields more complete and accurate segmentation results. These results validate PTC as a simple and effective approach to improving visual-text alignment.