arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过从细到粗的监督实现SigLIP-HD

SigLIP-HD by Fine-to-Coarse Supervision

Lihe Yang, Zhen Zhao, Hengshuang Zhao

arXiv 2607.09488首次发表:更新:

发表机构

The University of Hong Kong; Shanghai AI Laboratory(香港大学; 上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究如何在低成本下实现精细视觉感知,提出SigLIP-HD,采用从细到粗监督设计,基于SigLIP 2模型构建,在相同推理预算下能产生更好视觉令牌,在多基准测试中结果优于基线模型。

AI 中文摘要

高质量视觉表示是计算机视觉领域长期追求的目标。在多模态语言模型(MLLMs)中,输入更高分辨率图像可产生更细粒度的视觉令牌,但会增加计算和设计复杂性。本文研究如何在低成本且不使用大图像的情况下实现精细视觉感知。提出SigLIP-HD,核心是简单的从细到粗监督设计,基于SigLIP 2模型构建,使粗特征模仿高分辨率版本的细粒度特征。在相同推理预算下,该模型产生更好的视觉令牌,在多个MLLM基准测试中验证,结果优于基线模型,尤其在OCR相关任务上。

英文摘要

High-quality visual representation is a long-standing pursuit in computer vision. In the context of multimodal LLMs (MLLMs), feeding higher-resolution images can produce more fine-grained visual tokens. However, it introduces additional computational and design complexity, due to multiple forward passes and post-processing of increased tokens. Before simply adopting a higher resolution, have we truly unlocked the model's full perception capability at a standard resolution? Therefore, we study an interesting problem: how to achieve fine visual perception under lower cost without larger images. We present SigLIP-HD in this work. The core is a highly simple fine-to-coarse supervision design. We enforce the coarse feature of a mid-resolution image to mimic the fine-grained feature of its high-resolution version. We build this framework on the advanced SigLIP 2 model. Our final model produces better visual tokens at exactly the same inference budget. It is validated on extensive MLLM benchmarks and consistently delivers stronger results than our baseline model, especially on OCR-related tasks.

CommentsICLR 2026. Code and model: https://github.com/LiheYoung/SigLIP-HD

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑