arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05864cs.CV

用于域泛化分割的层级提示注入器

Hierarchical Prompt Injector for Domain Generalization Segmentation

Xin Kun Lin, Ruoyu Guo, Jiaqi Guo, Maurice Pagnucco, Yang Song

首次发表
浏览论文内容

中文总结 AI 辅助

针对域泛化语义分割,提出空间层级提示和层级提示注入器,实现空间自适应提示注入,在合成到真实和真实到真实基准上分别达到70.62%和72.74%的mIoU。

中文摘要 AI 辅助

域泛化语义分割(DGSS)是一项具有挑战性的任务,因为视觉模型通常依赖于跨域变化的低级外观线索。相比之下,结构属性表现出跨域稳定性,这促使在DGSS中使用结构先验。现有方法使用提示学习将这些先验迁移到DGSS模型中,但通常将每个类别编码为单一的全局提示。此外,这些方法将提示均匀应用于所有像素,当由于视角变化、遮挡和环境变化而仅可见对象区域的子集时,无法提供适应机制。我们通过提出空间层级提示(SHP)来解决这一问题,SHP为每个类别丰富了区域级几何锚点,从不同视角捕获结构外观,确保在任意视角下互补覆盖。此外,我们提出了层级提示注入器(HPI),它能够在基础模型中实现空间自适应的提示注入。HPI通过建模提示与视觉特征的语义相关性和空间影响来空间接地提示。考虑到学习空间和语义感知的提示注入的难度,我们进一步引入了辅助监督,以将层级提示与其对应的对象区域对齐。我们在合成到真实和真实到真实的基准上分别实现了70.62%和72.74%的mIoU。代码和检查点已在此https URL发布。

英文摘要

Domain Generalized Semantic Segmentation (DGSS) is a challenging task, as vision models often rely on low-level appearance cues that change across domains. In contrast, structural attributes exhibit cross-domain stability, motivating the use of structural priors for DGSS. Existing methods use prompt learning to transfer such priors into DGSS models, but typically encode each class as a single holistic prompt. Moreover, these methods apply prompts uniformly to all pixels, offering no mechanism to adapt when only a subset of object regions is visible due to viewpoint changes, occlusion, and environmental variation. We address this with \textbf{Spatial Hierarchical Prompts (SHP)} that enrich each class with region-level geometric anchors capturing structural appearance from distinct viewing angles, ensuring complementary coverage under arbitrary viewpoints. Additionally, we propose the \textbf{Hierarchical Prompt Injector (HPI)}, which enables spatially adaptive prompt injection in foundation models. HPI spatially grounds prompts by modeling their semantic relevance and spatial influence with visual features. Considering the difficulty of learning spatially and semantically aware prompt injection, we further introduce auxiliary supervision to align hierarchical prompts with their corresponding object regions. We achieve 70.62\% and 72.74\% mIoU on synthetic-to-real and real-to-real benchmarks, respectively. Code and checkpoints are released at https://github.com/MosukFate/HPI

发表机构

  • University of New South Wales(新南威尔士大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑