arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新思考视觉基础模型适应中的逐层信息分配

Rethinking Layer-Wise Information Allocation for Vision Foundation Model Adaptation

Yuqi Li, Xi Xiao, Yunbei Zhang, Lin Zhao, Yu Li, Aiden Zhao, Tianyang Wang, Hao Xu, Yingli Tian

arXiv 2607.21973首次发表:更新:

发表机构

City College of New York, CUNY; University of Alabama at Birmingham; Tulane University; Northeastern University; George Washington University; Capital One; Harvard University(纽约市立大学城市学院; 阿拉巴马大学伯明翰分校; 杜兰大学; 东北大学; 乔治华盛顿大学; 第一资本金融公司; 哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究视觉基础模型适应中的逐层信息分配问题,提出基于信息瓶颈原理的PIB框架,规范逐层压缩与充分性权衡,促进跨层信息连贯,实验表明该方法在多数据集上表现出色,能解释提示容量缩放行为等。

AI 中文摘要

视觉基础模型越来越多地被用作下游视觉识别的冻结主干,使参数高效适应成为核心问题。基于提示的适应方法(如视觉提示调整)提供了一种轻量级的模型专门化方式,但其逐层行为仍未被充分理解。我们认为这一限制不仅是优化问题,更是逐层信息分配问题。受信息瓶颈原理启发,我们引入了提示信息瓶颈(PIB)框架,该框架规范了逐层压缩-充分性权衡,并促进了更连贯的跨层信息路径。实验表明,PIB在34个数据集上表现出色,在FGVC上达到92.1%,在HTA上达到93.01%,在VTAB-1k上达到77.33%,平均仅调整0.35%的参数。PIB有助于解释提示容量缩放的非单调行为,减少对捷径的依赖,并提高在分布变化和细粒度识别设置下的鲁棒性。

英文摘要

Vision foundation models are increasingly reused as frozen backbones for downstream visual recognition, making parameter-efficient adaptation a central problem. Prompt-based adaptation, including Visual Prompt Tuning (VPT), provides a lightweight way to specialize these models, but its layer-wise behavior remains poorly understood: performance is sensitive to prompt depth, placement, and task distribution, and gains on standard in-domain benchmarks do not always translate into robust generalization. We argue that this limitation is not solely an optimization issue, but a layer-wise information allocation issue: existing prompt-based methods lack principled control over what prompt-conditioned representations should preserve, suppress, and propagate across depth. Inspired by the Information Bottleneck principle, we introduce Prompted Information Bottlenecks (PIB), a framework that regularizes layer-wise compression-sufficiency trade-offs and promotes a more coherent cross-layer information path. The key idea is that effective adaptation should be minimal yet sufficient, retaining task-relevant local evidence in earlier layers while progressively discarding nuisance factors and redundant details in deeper layers. Extensive experiments show that PIB achieves strong performance across 34 datasets, reaching 92.1% on FGVC, 93.01% on HTA, and 77.33% on VTAB-1k, while tuning only 0.35% parameters on average across the main settings. Beyond benchmark accuracy, PIB helps explain the non-monotonic behavior of prompt capacity scaling, reduces shortcut reliance, and improves robustness under distribution shift and fine-grained recognition settings. These results position PIB as both a practical method and an information-allocation perspective for adapting frozen vision foundation models. Our code is available at https://github.com/itsnotacie/MM-26-PIB

CommentsAccepted by ACM MM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑