arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TwinICL:通过配对反事实诊断多模态上下文学习

TwinICL: Diagnosing Multimodal In-Context Learning through Paired Counterfactuals

Zihan Xue, Po-Yi Lu, Serhii Honcharenko, Zih-Ching Chen, Hsuan-Tien Lin, Nanyun Peng, I-Hung Hsu, Kuan-Hao Huang

arXiv 2609.15028首次发表:更新:

发表机构

University of California, Los Angeles; National Taiwan University; Texas A&M University; NVIDIA AI Technology Center; Arena Intelligence Inc(加州大学洛杉矶分校; 国立台湾大学; 德克萨斯A&M大学; 英伟达AI技术中心; Arena Intelligence公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TwinICL提出配对反事实基准,发现多模态ICL弱于文本ICL,组合干预可恢复性能,并揭示模态差距及演示的双重作用。

AI 中文摘要

上下文学习(ICL)使模型能够从演示中推断任务,但现有基准通常缺乏用于跨模态比较ICL性能的匹配文本和图像版本。我们引入了TwinICL,一个程序化生成的基准,提供此类配对以进行受控比较。在六个开放权重模型和38个任务中,多模态ICL始终不如纯文本ICL,差距因任务族而异。为了测试这一差距是否可以恢复,我们针对视觉访问、任务框架和推理进行了三项干预。尽管个体效应有限或不一致,但它们的组合在诊断子集上恢复了强大的多模态ICL性能。为了区分执行任务的困难与推断任务的困难,我们用明确的任务指令评估模型,揭示即使在任务已知时也存在模态差距。然后,我们检查添加演示输入和输出如何重塑这一差距,强调演示作为待处理额外上下文和任务证据的双重作用。数据集可在https://this https URL获取。

英文摘要

In-context learning (ICL) enables models to infer tasks from demonstrations, but existing benchmarks generally lack matched text and image versions needed to compare ICL performance across modalities. We introduce TwinICL, a procedurally generated benchmark providing such pairs for controlled comparison. Across six open-weight models and 38 tasks, multimodal ICL consistently underperforms text-only ICL, with gaps varying by task family. To test whether this gap can be recovered, we target visual access, task framing, and reasoning through three interventions. Their combination recovers strong multimodal ICL performance on a diagnostic subset, despite limited or inconsistent individual effects. To distinguish difficulties in executing tasks from those in inferring them, we evaluate models with explicit task instructions, revealing a modality gap even when the task is known. We then examine how adding demonstration inputs and outputs reshapes this gap, highlighting demonstrations' dual role as additional context to process and evidence about the task. The dataset is available at https://github.com/lab-flair/TwinICL.

Comments21 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑