arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12724cs.LG

MAG:流形引导的半监督多模态上下文学习

MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning

Zirui Cheng, Xun Xu, Tiankai Chen, Fady Rezk, Bowen Zheng, Xiaodong Shi, Shijie Li, Kangkang Lu, Bharadwaj Veeravalli, Nancy F. Chen

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出MAG框架,通过将未标记多模态数据用于半监督传播的演示选择,提升多模态大语言模型的上下文学习性能,在8个基准的标签稀缺场景下优于强基线。

中文摘要 AI 辅助

多模态大语言模型(MLLM)的少样本上下文学习(ICL)无需参数更新即可实现任务适配,但其性能高度依赖所选演示样本的质量与覆盖范围。尽管未标记多模态数据十分丰富,如何将其用于ICL仍不明确。本文提出MAG(流形引导半监督上下文演示选择),这一利用未标记数据提升多模态ICL的高效框架。MAG将演示选择建模为多模态图上的半监督传播问题,采用两阶段策略:(i)相关度得分传播识别出一组紧凑的高影响力未标记样本以生成伪标签,降低MLLM推理成本;(ii)多模态相关度用于选择最终演示样本。研究发现,文本表示对相关度传播更有效,而视觉与文本模态对高质量演示选择均至关重要。在8个多模态基准上的实验表明,MAG在标签稀缺场景下始终优于强基线,且在有限伪标签预算下实现了显著提升。

英文摘要

Few-shot in-context learning (ICL) with multi-modal large language models (MLLMs) enables task adaptation without parameter updates, but its performance is highly sensitive to the quality and coverage of the selected demonstrations. While unlabeled multi-modal data is abundant, it remains elusive how to exploit them for ICL. We propose MAG (MAnifold-Guided semi-supervised in-context demonstra- tion selection), an efficient framework that leverages unlabeled data to improve multi-modal ICL. MAG formulates demonstration selection as a semi-supervised propagation problem on a multi-modal graph and adopts a two-stage strategy: (i) relevance score propagation identifies a compact set of high-impact unlabeled samples for pseudo-labeling, reducing MLLM inference cost; (ii) multi-modal relevance is used to select the final demonstrations. We show that textual represen- tations are more effective for relevance propagation, while both visual and textual modalities are crucial for high-quality demonstration selection. Experiments on eight multi-modal benchmarks demonstrate that MAG consistently outperforms strong baselines in label-scarce regimes, achieving significant gains with a limited pseudo-labeling budget.

↑