发表机构
University of Wisconsin–Madison; Northwestern University(威斯康星大学麦迪逊分校; 西北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
STAVE用两个任务特定向量取代上下文演示,实现零射击推理成本,在多模态和文本任务上以更少参数达到或超越最先进方法。
AI 中文摘要
上下文学习(ICL)通过少量演示(demos)将冻结的大型多模态模型(LMMs)适配到新任务,但在每次查询时都会重新编码这些演示,其中每个演示图像最多增加数百个视觉令牌。无演示方法通过紧凑的任务状态消除了这一成本。然而,它们将任务状态添加到按任务搜索的位置或每个解码器层,导致任务参数随深度增长。此外,插入的令牌或键无法改变原始提示在层内分配注意力的方式。为解决这些问题,我们提出了基于嵌入的结构化任务适配(STAVE),该方法用两个任务特定向量取代演示,并将它们添加到现有输入嵌入中。具体而言,一个读出向量更新生成答案的令牌,一个上下文向量更新其他结构令牌组。两者均在包含和不包含演示的提示上使用答案标签进行训练。我们通过损失的一阶分析和边际界从理论上证明了这些设计选择的合理性。在六个LMM和五个大型语言模型上的大量实验表明,STAVE在多模态任务上以远更少的任务参数达到或超越最先进方法,并在18个文本任务上超越15次射击ICL和先前的任务向量,且推理成本为零射击。
英文摘要
In-context learning (ICL) adapts frozen large multimodal models (LMMs) to new tasks from a few demonstrations (demos), but re-encodes them at every query, where each demo image adds up to hundreds of visual tokens. Demo-free methods remove this cost with a compact task state. However, they add it at locations searched per task or at every decoder layer, where task parameters grow with depth. Moreover, inserted tokens or keys cannot change how the original prompt divides its attention within a layer. To address these issues, we propose Structured Task Adaptation via Embeddings (STAVE), which replaces demos with two task-specific vectors added to existing input embeddings. Specifically, a readout vector updates the answer-producing tokens and a context vector updates the other structural token groups. Both are trained with answer labels on prompts with and without demos. We justify these design choices theoretically using a first-order analysis of the loss and a margin bound. Extensive experiments on six LMMs and five large language models show that STAVE matches or outperforms state-of-the-art methods on multimodal tasks with far fewer task parameters and surpasses 15-shot ICL and prior task vectors on 18 text tasks, all at zero-shot inference cost.
CommentsTechnical report