发表机构
Gachon University(嘉泉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SPACE-CLIPv2提出在冻结CLIP标记空间中进行固定局部邻域聚合,通过门控残差更新和高通分支,在NYU Depth V2和零样本iBims-1上提升单目深度估计的边界与平面几何精度。
AI 中文摘要
视觉-语言基础模型如CLIP提供了强大的语义表示,但其patch标记并未直接针对密集度量几何进行优化。SPACE-CLIP表明,冻结的CLIP特征可通过层组特征融合支持单目深度估计,但仍未解决如何组合相邻CLIP标记以恢复精细局部结构的问题。我们提出SPACE-CLIPv2,一种冻结骨干的深度解码器,在CLIP标记空间中聚合固定的局部邻域。在选定的解码器阶段,模型采样固定的标记模板,预测聚合权重,并通过门控残差更新注入所得响应。一个标记空间高通分支进一步保留浅层局部对比度。在NYU Depth V2上,SPACE-CLIPv2相比匹配的SPACE-CLIP基线有所改进,而五次种子实验一致支持固定而非学习偏移的采样。零样本iBims-1评估进一步改善了边界和平面几何度量。这些结果支持受约束的局部标记聚合作为从冻结CLIP表示解码几何的实用机制。
英文摘要
Vision-language foundation models such as CLIP provide strong semantic representations, but their patch tokens are not directly optimized for dense metric geometry. SPACE-CLIP showed that frozen CLIP features can support monocular depth estimation through layer-group feature fusion, yet it leaves open how neighboring CLIP tokens should be combined to recover fine local structure. We present SPACE-CLIPv2, a frozen-backbone depth decoder that aggregates fixed local neighborhoods in CLIP token space. At selected decoder stages, the model samples a fixed token stencil, predicts aggregation weights, and injects the resulting response through a gated residual update. A token-space high-pass branch further preserves shallow local contrast. On NYU Depth V2, SPACE-CLIPv2 improves over a matched SPACE-CLIP baseline, while five-seed experiments consistently favor fixed over learned-offset sampling. Zero-shot iBims-1 evaluation further improves boundary and planar-geometry measures. These results support constrained local token aggregation as a practical mechanism for decoding geometry from frozen CLIP representations.
Journal ref2026 Asian Conference on Computer Vision (ACCV)