arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SeCo-SBIR:用于零样本草图-图像检索的语义一致提示学习

SeCo-SBIR: Semantically Consistent Prompt Learning for Zero-Shot Sketch-Based Image Retrieval

Long Hoang Dang, Tuan Nguyen Huu, Nguyen Minh Hieu, Tu Minh Phuong

arXiv 2608.03120首次发表:更新:

AI 中文总结

SeCo-SBIR是解决CLIP适配ZS-SBIR时域适配与泛化矛盾的语义一致提示学习框架,结合多模态提示、一致性约束与多目标损失,在三类ZS-SBIR基准上达最优性能。

AI 中文摘要

将CLIP通过提示学习适配零样本草图-图像检索(ZS-SBIR)面临一个核心矛盾:模型需通过任务特定的适配缩小草图与照片的域差距,但增加的灵活性可能导致模型过拟合已见过的训练类别,削弱CLIP的零样本泛化能力。本文提出SeCo-SBIR,一种语义一致的提示学习框架,从两方面解决该矛盾:第一,文本引导的多模态提示策略将可学习提示向量输入CLIP的文本编码器,再通过可学习耦合函数将中间表示投影到视觉编码器的每一层;由于文本编码器已通过大规模语言监督学习到鲁棒、抽象的类别级语义,该机制将可迁移的语义知识直接注入视觉通路,使模型适配草图-照片域的同时,天然偏向对未见类别的泛化。第二,基于扰动的一致性约束通过非对称InfoNCE目标,将适配后的模型与冻结的CLIP参考分支对齐,解决可学习耦合函数带来的剩余过拟合风险——增强输入送入冻结分支,干净输入送入可训练分支,将学习到的表示锚定到CLIP的可泛化特征空间。结合轻量适配器与包含三元组损失、NT-Xent损失和分类项的多目标损失,SeCo-SBIR在类别、广义及跨数据集设置的三个标准ZS-SBIR基准上均取得了最优结果。

英文摘要

Adapting CLIP for zero-shot sketch-based image retrieval (ZS-SBIR) via prompt learning faces a fundamental tension: the model must bridge the sketch-photo domain gap through task-specific adaptation, yet the added flexibility risks overfitting to seen training categories and eroding CLIP's zero-shot generalization. We present SeCo-SBIR, a semantically consistent prompt learning framework that resolves this tension from both sides. First, a text-guided multi-modal prompting strategy routes learnable prompt vectors through CLIP's text encoder and projects the resulting intermediate representations into the visual encoder at every layer via learnable coupling functions. Because the text encoder has already learned robust, abstract category-level semantics from large-scale language supervision, this mechanism injects transferable semantic knowledge directly into the visual pathway - adapting the model to the sketch-photo domain while inherently favoring generalization to unseen classes. Second, a perturbation-based consistency constraint addresses the residual overfitting risk from the learnable coupling functions by aligning the adapted model with a frozen CLIP reference branch using an asymmetric InfoNCE objective - augmented inputs feed the frozen branch while clean inputs feed the trainable branch - anchoring the learned representations to CLIP's generalizable feature space. Together with lightweight adapters and a multi-objective loss combining triplet, NT-Xent, and classification terms, SeCo-SBIR achieves state-of-the-art results on all three standard ZS-SBIR benchmarks across categorical, generalized, and across-dataset settings.

CommentsACMMM 26

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑