arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07729cs.CV

Foveated Compression:面向令牌高效视觉语言模型的选择性高分辨率保留

Foveated Compression: Selective High-Resolution Preservation for Token-Efficient VLMs

Donghyun Han, Jangho Park, Yuseok Bae

首次发表
浏览论文内容

中文总结 AI 辅助

针对视觉语言模型令牌预算固定问题,提出Foveated Compression方法,通过混合分辨率令牌和轻量选择器保留局部高保真,实验显示局部保真并非普遍最优,区域选择与压缩质量存在互补瓶颈。

中文摘要 AI 辅助

视觉令牌是视觉语言模型推理成本的主要来源,然而简单的图像下采样仍然是一个出人意料的强大压缩基线。这引出了一个互补性问题:在固定的令牌预算下,视觉保真度应保留在何处?我们引入了Foveated Compression(中央凹压缩),该方法对全分辨率图像编码一次,并用原生分辨率和压缩分辨率视觉令牌的混合来表示它。一个行为自蒸馏的Foveated Merger(中央凹合并器)压缩局部视觉令牌,同时保持与其原生对应物的兼容性,而一个轻量级的Foveated Selector(中央凹选择器)使用穷举的预算匹配干预监督,从九个空间单元中选择一个以原生分辨率保留。在11.11%视觉令牌下,均匀Foveated Compression与等令牌下采样相比无显著配对差异。在20.99%下,学习的选择器显著优于随机和固定分配,但仍低于强整体图像缩放,表明局部保真度并非普遍更优。一个预算匹配的区域选择预言机达到82.73宏平均准确率,而学习的选择器为69.61,揭示了在同一空间动作空间内的巨大提升空间。匹配的探测进一步表明,预测压缩何时破坏答案的信号在语言模型计算后比轻量级无预填充选择器更容易获取。这些结果揭示了区域选择和压缩区域保真度中的互补瓶颈。

英文摘要

Visual tokens are a major source of inference cost in vision-language models, yet simple image downsampling remains a surprisingly strong compression baseline. This raises a complementary question: under a fixed token budget, where should visual fidelity be preserved? We introduce Foveated Compression, which encodes a full-resolution image once and represents it with a mixture of native- and compressed-resolution visual tokens. A behaviorally self-distilled Foveated Merger compresses local visual tokens while preserving compatibility with their native counterparts, and a lightweight Foveated Selector chooses one of nine spatial cells to retain at native resolution using exhaustive budget-matched intervention supervision. At 11.11% visual tokens, uniform Foveated Compression shows no significant paired difference from iso-token downsampling. At 20.99%, the learned selector significantly outperforms random and fixed allocation, but remains below strong whole-image resizing, showing that localized fidelity is not universally preferable. A budget-matched region-choice oracle reaches 82.73 macro accuracy versus 69.61 for the learned selector, revealing substantial headroom within the same spatial action space. Matched probing further shows that signals predicting when compression breaks the answer are substantially more accessible after language-model computation than to the lightweight prefill-free selector. These results expose complementary bottlenecks in region selection and compressed-region fidelity.

发表机构

  • ETRI(韩国电子通信研究院)
  • Kyung Hee University(庆熙大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑