发表机构
Brown University(布朗大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出选择-实现假说,通过受控多模态任务和VQA基准研究发现,静态任务向量的成功取决于演示诱导变化的跨查询共享程度,为隐式多模态上下文学习的方法选择提供了统一经验理论。
AI 中文摘要
隐式多模态上下文学习(Implicit multimodal in-context learning)将演示压缩为内部干预,范围从静态任务向量(static task vectors)到查询条件变换(query-conditioned transformations)及注意力路由(attention routing)。尽管这些方法目标一致,但干预对查询的依赖方式、对模型的修改位置存在显著差异,导致对于给定任务,尚不清楚需要多大的额外复杂度。本文提出选择-实现假说(Selection--Realization Hypothesis),该假说将演示视为诱导出一组紧凑的内部变化,查询从中进行选择,而模型的计算则约束所选变化的实现方式。我们使用受控多模态任务评估该假说,此类任务中查询依赖程度变化,但基础任务原语(task primitives)或提示格式保持不变。通过对比正确演示与匹配的反事实(counterfactuals),我们测量显式M-ICL的结构并测试其是否能预测干预行为。研究发现,静态任务向量的成功与演示诱导的变化在多大程度上跨查询共享密切相关;当显式M-ICL包含局部加性偏移无法恢复的查询特定或分布式结构时,额外的干预复杂度才会发挥作用。这些关系可扩展至自然视觉问答(VQA)基准,支持在无法获取测试性能时进行成本感知的方法选择。我们的结果为何时可将演示压缩为任务向量、何时需要更具表达力的干预提供了统一的经验理论。
英文摘要
Implicit multimodal in-context learning compresses demonstrations into internal interventions, ranging from static task vectors to query-conditioned transformations and attention routing. Despite their common goal, these methods differ substantially in how the intervention depends on the query and where it modifies the model, leaving unclear which additional complexity is necessary for a given task. We propose the Selection--Realization Hypothesis. It views demonstrations as inducing a compact family of internal changes from which the query selects, while the model's computation constrains how the selected change can be implemented. We evaluate this account using controlled multimodal tasks in which query dependence varies without changing the underlying task primitives or prompt format. By contrasting correct demonstrations with matched counterfactuals, we measure the structure of explicit M-ICL and test whether it predicts intervention behavior. We find that the success of a static task vector is closely tied to how much of the demonstration-induced change is shared across queries. Additional intervention complexity becomes useful when explicit M-ICL contains query-specific or distributed structure that a local additive shift cannot recover. These relationships extend to natural VQA benchmarks and support cost-aware method selection without access to test performance. Our results provide a unified empirical theory of when demonstrations can be compressed into a task vector and when a more expressive intervention is warranted.
CommentsAccepted by Empirical Theory in Representation Learning @ ECCV 2026, Oral