arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05627cs.CV

SCI-CLIP:基于参考记忆的无训练开放词汇分割的以段为中心推理框架

SCI-CLIP: Segment-Centric Inference with Reference Memory for Training-Free Open-Vocabulary Segmentation

Mohamad Zamini, Diksha Shukla

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出SCI-CLIP框架,以段为中心推理,无需训练即可将CLIP特征转化为高质量密集预测,在八个基准测试中提升开放词汇分割性能。

中文摘要 AI 辅助

无训练开放词汇分割受限于缺失的推理抽象。冻结的视觉-语言特征以patch级别生成,但密集预测需要一个单元,同时控制特征交互、空间支持、上下文恢复和基于检索的校正。我们提出SCI-CLIP,一个以段为中心的推理框架,其核心原则是相同的区域抽象应组织密集开放词汇预测的所有阶段。SCI-CLIP首先在冻结的视觉token上生成区域一致的交互图,随后通过在该图上传播值来重建密集特征,仅在局部证据不足时用选择性跨窗口支持增强这些特征。接着,相同的段抽象被用于构建和查询离线参考记忆,使示例检索与预测所基于的单元对齐。SCI-CLIP无需任何训练即可将冻结的CLIP风格特征转化为空间连贯、上下文感知且兼容检索的密集预测。SCI-CLIP持续提升密集预测的结构质量、上下文推理的鲁棒性以及基于示例校正的对齐度,在八个基准测试中实现了更强的开放词汇分割性能。项目代码可在:this https URL获取。

英文摘要

Training-free open-vocabulary segmentation remains limited by a missing inference abstraction. Frozen vision-language features are produced at patch level, yet dense prediction requires a unit that simultaneously governs feature interaction, spatial support, contextual recovery, and retrieval-based correction. We present SCI-CLIP, a segment-centric inference framework built around the principle that the same region abstraction should organize all stages of dense open-vocabulary prediction. SCI-CLIP first induces a region-consistent interaction graph over frozen visual tokens, then reconstructs dense features by propagating values over this graph, augmenting them with selective cross-window support only where local evidence is insufficient. The same segment abstraction is subsequently used to construct and query an offline reference memory, aligning exemplar retrieval with the units on which prediction is made. SCI-CLIP turns frozen CLIP-style features into spatially coherent, context-aware, and retrieval-compatible dense predictions without any training. SCI-CLIP consistently improves the structural quality of dense predictions, the robustness of contextual reasoning, and the alignment of exemplar-based correction, yielding stronger open-vocabulary segmentation across eight benchmarks. Project code is available at: https://github.com/mzamini92/SCICLIP.

↑