arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉语言模型中的上下文坍缩及其缓解方法

In-Context Collapse in Vision-Language Models and How to Mitigate it?

Mohammad Rostami

arXiv 2608.02830首次发表:更新:

发表机构

Amazon Generative AI Innovation Center(亚马逊生成式AI创新中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文发现视觉语言模型存在随演示增多的上下文坍缩问题,通过因果定位将其归因于视觉-语言整合通路,提出CircA干预方法可有效缓解该问题并实现抗坍缩能力迁移。

AI 中文摘要

多示例上下文学习(ICL)允许视觉语言模型(VLM)无需更新权重即可从图像-标签演示中自适应,且通常认为提供的演示越多,性能越好。但本文显示相反情况:随着演示积累,部分VLM会发生“上下文坍缩”,这是一种急剧、有时是灾难性的准确率下降,涉及合成分类、自然图像分类和VQA基准测试,部分模型的准确率甚至低于随机水平,而输出仍保持正常。在0.5B至11B的开源VLM面板以及前沿模型Claude Sonnet 4.5中,该坍缩呈分级特征。研究发现两种能力可分离:对积累演示的鲁棒性和上下文学习新规则的能力,二者组合产生三种可复现的状态。参数匹配的损伤与修复实验将坍缩因果定位到视觉-语言整合通路:在连接器及早期/中间层添加适配器可恢复真实学习(16示例时重映射准确率从0.39升至0.91),而在晚期读出层添加同等容量适配器则无效。本文提出CircA,其核心是一次性整合“疫苗”:在一项合成任务上训练一次后,可将抗坍缩能力迁移至未见任务族(在CIFAR、Fashion上从随机水平分别升至0.71、0.60)。此外,上下文整合的最佳层并非基于权重的整合的最佳层,晚期读出层在参数更少时实现更高准确率和更少遗忘。该坍缩是视觉-语言接口的整合失败,可通过轻量、可迁移的干预措施纠正。

英文摘要

Many-shot in-context learning (ICL) lets vision-language models (VLMs) adapt from image--label demonstrations without weight updates, and is widely assumed to improve as more demonstrations are supplied. We show the opposite: as demonstrations accumulate, a subset of VLMs undergo an \emph{in-context collapse}, a sharp, sometimes catastrophic accuracy drop spanning synthetic classification, natural-image classification, and VQA benchmarks, in some models falling below chance while outputs remain well-formed. Across an open VLM panel ($0.5$B--$11$B) and a frontier model (Claude Sonnet 4.5), the collapse is graded. Two capabilities turn out to be dissociable: robustness to accumulating demonstrations and the ability to learn a novel rule in context, their combinations yield three reproducible regimes. A parameter-matched lesion-and-rescue causally localizes the collapse to the vision-language integration pathway: an adapter on the connector and early/mid layers restores genuine learning (remap accuracy $0.39!\rightarrow!0.91$ at 16 shots), while an equal-capacity adapter on the late readout does not. We propose \textsc{CircA}, whose core is a one-time integration vaccine: trained once on one synthetic task, it transfers collapse-resistance to unseen task families (chance$\rightarrow$$0.71$/$0.60$ on CIFAR/Fashion). The layers best for in-context integration are not the layers best for weight-based consolidation, the late readout achieves higher accuracy and less forgetting at fewer parameters. The collapse is an integration failure at the vision--language interface, correctable by a lightweight, transferable intervention.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑