arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

展示示例:从图像集合中推断视觉概念

Show Me Examples: Inferring Visual Concepts from Image Sets

Nick Stracke, Kolja Bauer, Stefan Andreas Baumann, Miguel Angel Bautista, Josh Susskind, Björn Ommer

arXiv 2607.02402首次发表:更新:

发表机构

Munich Center for Machine Learning; Apple(慕尼黑机器学习中心; 苹果公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出VICIS任务评估视觉语言模型从图像集合中推断共享概念的能力,并设计训练框架与架构,使模型能提取概念嵌入并生成与查询一致的新图像。

AI 中文摘要

视觉语言模型(VLM)能够遵循复杂的文本指令,但难以从纯视觉上下文中进行推理。特别是,当前模型无法从示例图像集合中推断共享概念并将其应用于新输入。我们引入了从集合中推断视觉概念(VICIS)任务来评估这一能力。给定一个共享概念的少量图像上下文集和一张查询图像,模型必须生成保留上下文定义概念且与查询一致的新图像。我们发现,最先进的VLM在此任务上表现不佳,常常忽略视觉上下文或默认生成有偏见的图像。为解决这一差距,我们提出了一种训练框架和架构,学习从图像集合中推断视觉概念,并从查询中提取特定于概念的嵌入。在合成数据和大规模ImageNet/WordNet数据上的实验表明,我们的模型生成更准确和多样化的输出,并能泛化到未见概念和模态(如草图)。

英文摘要

Vision-language models (VLMs) can follow complex textual instructions, yet they struggle to reason from purely visual context. In particular, current models fail to infer shared concepts from sets of example images and apply them to new inputs. We introduce Visual Concept Inference from Sets (VICIS), a task that evaluates this capability. Given a small context set of images sharing a concept and a query image, the model must generate new images that preserve the context-defined concept while remaining consistent with the query. We show that state-of-the-art VLMs perform poorly on this task, often ignoring the visual context or defaulting to biased generations. To address this gap, we propose a training framework and architecture that learn to infer visual concepts from image sets and extract concept-specific embeddings from queries. Experiments on synthetic data and large-scale ImageNet/WordNet data show that our model generates more accurate and diverse outputs and generalizes to unseen concepts and modalities such as sketches.

Commentsfor code, view https://github.com/CompVis/set-learner

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑