发表机构
LIGM; CNRS; Univ Gustave Eiffel; ENPC; Institut Polytechnique de Paris; UC Berkeley(LIGM机构; 法国国家科学研究中心; 古斯塔夫·埃菲尔大学; 法国国立路桥学校; 巴黎综合理工学院; 加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种自底向上的单图像重复元素发现方法,通过重建目标学习可调原型,无需标注或分割,在FSC-147上验证了其有效性与可解释性。
AI 中文摘要
我们研究了从单张图像中发现重复元素的问题。与依赖大型标注数据集、精心整理的多图像集合或对象分割掩码的现有方法不同,我们证明单张图像足以以完全自底向上的方式学习有意义的对象模型,除了粗略的尺度先验外,无需任何先验知识。我们的方法通过重建目标学习重复元素的可调图像空间原型,使模型能够识别并合成同一图像内一致的物体实例。在FSC-147数据集的116张真实图像上的实验表明,我们的方法成功学习了连贯的元素模型,并在具有挑战性的图像上捕捉了类别内变化。定性结果显示,与经典分解、联合对齐和3D物体建模方法相比,我们的方法具有更优的重建效果和可解释的分解,同时保持了简单的2D公式。这些结果表明,有意义的物体发现可以仅从单图像学习中涌现。
英文摘要
We address the problem of discovering repeated elements from a single image. In contrast to existing approaches that depend on large annotated datasets, curated multi-image collections, or object segmentation masks, we show that a single image can suffice to learn a meaningful object model in a completely bottom-up fashion, without any prior knowledge beyond a coarse scale prior. Our method learns a tunable image-space prototype of the repeated elements through a reconstruction objective, enabling the model to identify and synthesize consistent object instances within the same image. Experiments on 116 real images from the FSC-147 dataset demonstrate that our method successfully learns coherent element models and captures intra-category variation on challenging images. Qualitative results reveal superior reconstructions and interpretable decompositions compared to classical decomposition, joint alignment, and 3D object modeling methods, while maintaining a simple 2D formulation. These results suggest that meaningful object discovery can emerge from single image learning alone.
CommentsAccepted to ECCV 2026. Project page: https://vayvi.github.io/repeated-elements/