发表机构
Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ENCORE是一个熵引导的视觉-语言理解框架,通过基于熵的裁剪策略和熵正则化训练,仅微调0.14%参数,在10个VQA基准上实现平均1.43%的准确率提升,达到2B参数VLMs的SOTA性能。
AI 中文摘要
视觉-语言模型(VLMs)在各类视觉-语言任务中表现良好,但基于Transformer的视觉编码器将图像划分为固定分辨率的子图像,这会损害轻量级VLMs中的对象完整性。现有方法仅关注视觉模态,未能动态保留与提示相关区域的完整性,从而限制了性能。在本研究中,我们观察到跨模态注意力的早期图像-文本熵与答案定位质量和任务准确性密切相关。基于这一发现,我们提出了ENCORE,这是一个包含两个组件的熵引导框架:在推理阶段,一种基于熵的裁剪策略(ECS)会评估一小部分候选裁剪区域,并选择熵最小的那个,以保留与提示相关的连续区域;在训练阶段,熵正则化训练(ERT)会为下一个词元预测添加一个熵项,该熵项会增强对关键视觉词元的注意力,同时降低对不相关词元的权重。在十个视觉问答(VQA)基准上的实验表明,仅微调0.14%参数的ENCORE,平均准确率提升了1.43%,并且在近期20亿参数的VLMs中达到了最先进的性能。我们的代码已在此URL发布。
英文摘要
Vision-Language Models (VLMs) perform well on diverse vision-language tasks, but transformer-based visual encoders split images into fixed-resolution sub-images, compromising object integrity in lightweight VLMs. Existing methods only focus on the visual modality and fail to dynamically preserve the integrity of prompt-relevant regions, limiting performance. In this work, we observe that the early-layer image-text entropy of cross-modal attention strongly correlates with answer grounding quality and task accuracy. Building on this finding, we propose \textbf{ENCORE}, an entropy-guided framework with two components: At inference, an \textbf{Entropy-based Cropping Strategy} (ECS) evaluates a small set of candidate crops and selects the one with minimal entropy, preserving contiguous regions relevant to the prompt. At training, \textbf{Entropy Regularization Training} (ERT) augments next-token prediction with an entropy term that sharpens attention on key visual tokens while down-weighting irrelevant ones. Experiments on ten VQA benchmarks show that ENCORE, fine-tuning only 0.14\% of parameters, achieves an average 1.43\% accuracy gain and state-of-the-art performance among recent 2B-parameter VLMs. Our code is released in https://github.com/baokou-fw2/ENCORE.
Journal refICASSP 2026
DOI:10.1109/ICASSP55912.2026.11463726