arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04392cs.CV

FAVE:用于高效细粒度视觉理解的凹形自适应视觉编码

FAVE: Foveated Adaptive Visual Encoding for Efficient Fine-Grained Visual Understanding

  • Elmore Family School of Electrical and Computer Engineering(埃尔莫尔家族电气与计算机工程学院)
  • Purdue University(普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

Amitangshu Mukherjee, Kaushik Roy

AI总结:

FAVE是一种轻量级可变分辨率ViT,通过选择性编码高敏锐度局部区域,在ImageNet、TextVQA、GQA等任务上以更低计算量提升细粒度视觉与多模态理解性能

AI中文摘要:

细粒度视觉理解依赖于局部细节,但视觉编码器面临着高成本的全图高分辨率处理与紧凑的全局编码之间的权衡,后者会削弱这类证据。受人类主动视觉的启发,我们将“注视何处”与“编码何物”分离开,聚焦于后者并引入FAVE(Foveated Adaptive Visual Encoding,凹形自适应视觉编码),这是一种轻量级可变分辨率ViT,能以高敏锐度编码外部选定区域,同时保留原生几何结构。我们首先在受控的小目标场景中使用Oracle真值裁剪来单独研究该编码问题:在最大原生边长为96像素的ImageNet目标上,FAVE在相同裁剪窗口上相比固定分辨率ViT,Top-1准确率提升9.4个百分点,同时FLOPs降低12.7倍;提高全局分辨率或骨干网络容量无法恢复相同性能。随后,我们将FAVE作为补充局部分支集成到FastVLM中,其局部标记与FastVLM的全局视觉标记结合,原始全局通路和语言模型保持冻结状态:仅使用最多16个额外局部标记时,FAVE在TextVQA任务上提升1.60个百分点,相比SmolVLM2-2.2B实现3.3倍的受控TTFT加速;在GQA属性问题上,它提升FastVLM-1.5B 1.31个百分点,将优势扩展至文本之外,同时缩小与FastVLM-7B的差距。这些结果共同表明,选择性分配高敏锐度局部容量,可为细粒度理解小目标、文本和属性提供高效的全局表示与模型缩放的补充。

英文摘要:

Fine-grained visual understanding depends on local detail, yet visual encoders face a trade-off between costly full-image high-resolution processing and compact global encoding that can weaken such evidence. Inspired by human active vision, we separate where to look from what to encode. We focus on the latter and introduce FAVE (Foveated Adaptive Visual Encoding), a lightweight variable-resolution ViT that encodes externally selected regions at high acuity while preserving native geometry. We first isolate this encoding problem using oracle ground-truth crops in a controlled small-object regime. On ImageNet objects with a native maximum side of 96 pixels, FAVE improves Top-1 by 9.4 points over a fixed-resolution ViT on the same crop window with 12.7 times lower FLOPs. Increasing global resolution or backbone capacity does not recover the same operating point. We then integrate FAVE as a complementary local branch in FastVLM. Its local tokens are combined with FastVLM's global visual tokens, while the original global pathway and language model remain frozen. With at most 16 additional local tokens, FAVE improves TextVQA by 1.60 points and achieves a 3.3 times controlled TTFT speedup over SmolVLM2-2.2B. On GQA attribute questions, it improves FastVLM-1.5B by 1.31 points, extending the benefit beyond text while narrowing the gap to FastVLM-7B. Together, these results show that selectively allocating high-acuity local capacity provides an efficient complement to broader global representations and model scaling for fine-grained understanding of small objects, text, and attributes.

补充信息

↑