发表机构
Eightfold AI; Inception42; Indian Institute of Technology Kharagpur(Eightfold AI; Inception42; 印度理工学院卡哈尔普尔分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对VLM细粒度视觉感知不足,提出融合DINOv3与CleanDIFT并文本对齐的TDDN模型,以少量对齐数据在检索和分割上超越CLIP,并引入Puzzle Perception数据集验证其优势。
AI 中文摘要
结构化视觉推理(如图像拼图)需要细粒度的视觉感知,而当前视觉语言模型(VLM)缺乏这种能力。基于CLIP的ViT骨干构建的VLM以细粒度细节换取高层语义,我们证明这种损失会向下游传播。为恢复该能力,我们将DINOv3和CleanDIFT表示融合到感知编码器(DiffusedDINO)中,并将其与RoBERTa-L对齐,得到文本对齐模型TDDN,从而保留这一感知优势:在冻结骨干且仅使用约59万对齐样本的情况下,TDDN在图像-文本检索上匹配CLIP,并在四项设置中的三项上超越它。同时,TDDN的密集预测准确率是CLIP的三倍多(ADE20K从5.20提升至18.11 mIoU,COCO-Stuff从7.35提升至24.44),尽管CLIP使用了大规模训练语料。在通用对比编码器中,TDDN在分割基准上领先,包括SigLIP 2。我们进一步引入Puzzle Perception,一个用于探测细粒度空间理解的分割和视觉问答数据集,在该数据集上TDDN的分割准确率是CLIP的两倍(从11.04提升至22.51 mIoU)。
英文摘要
Structured visual reasoning, such as image puzzles, demands fine-grained visual perception, an ability current Vision Language Models (VLMs) lack. VLMs built on CLIP-based ViT backbones trade fine-grained detail for high-level semantics, and we show this loss propagates downstream. To recover it, we fuse DINOv3 and CleanDIFT representations into a perception encoder (DiffusedDINO) and align it with RoBERTa-L, yielding a text-aligned model TDDN that preserves this perceptual advantage: with frozen backbones and only $\sim$590K alignment pairs, TDDN matches CLIP on image-text retrieval, surpassing it on three of four settings. It does so while more than tripling CLIP's dense-prediction accuracy (ADE20K 5.20 $\to$ 18.11 mIoU, COCO-Stuff 7.35 $\to$ 24.44), despite CLIP's massive training corpus. TDDN leads on segmentation benchmarks among general-purpose contrastive encoders, including SigLIP$\,$2. We further introduce Puzzle Perception, a segmentation and visual question answering dataset that probes fine-grained spatial understanding, on which TDDN doubles CLIP's segmentation accuracy (11.04 $\to$ 22.51 mIoU).