arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LinSlot:利用线性表示假设从基于槽的对象表示中进行无监督属性发现

LinSlot: Exploiting Linear Representation hypothesis for unsupervised attribute discovery from slot based object representation

Sanket Gandhi, Utkarsh Giri, Varun Subramanium, Rohan Paul, Parag Singla

arXiv 2610.10722首次发表:更新:

发表机构

Yardi School of AI, IIT Delhi; Department of Computer Science and Engineering, IIT Delhi(德里印度理工学院亚迪人工智能学院; 德里印度理工学院计算机科学与工程系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对从原始图像学习对象与属性解耦表示的问题,提出融入线性表示假设(LRH)的LinSlot架构,联合发现对象和属性表示,在多数据集上提升了DCI分数,还可用于图像编辑。

AI 中文摘要

本文研究从原始非结构化图像数据中学习对象及其属性的解耦表示的问题。基于槽的方法在从图像中无监督学习对象表示方面已取得显著成功。基于块槽注意力的方法通过假设对象表示可分解为属性的均匀分解,将该框架扩展到属性表示,但这种假设可能并非最优,从而限制了所学表示的质量。因此,我们研究一种联合发现对象和属性表示的框架。我们的核心贡献是利用线性表示假设(Linear Representation Hypothesis, LRH),该假设提出可组合概念可表示为槽表示中的线性加性子空间。基于这一见解,我们提出了一种连接图像、槽(对象)和块(属性)的概率模型。我们提出了一种架构,该架构利用块注意力将属性表示连接到槽,并在对象和属性表示空间中都融入了LRH。该架构有效优化了所提出图形模型的证据下界(Evidence Lower Bound, ELBO)。我们的实验表明:(i)能有效发现解耦的对象和属性表示;(ii)提供了槽空间中LRH的经验证据;(iii)由于所学表示的解耦性和可解释性,具备执行图像编辑的能力。我们在多个数据集上的实验表明,与最先进的方法相比,DCI分数有所提升。

英文摘要

This paper studies the problem of learning disentangled representations of objects and their attributes from raw, unstructured image data. Slot-based methods have shown considerable success in unsupervised learning of object representations from images. Block-slot attention-based methods extend this framework to attribute representations by assuming a uniform factorization of object representations into attributes, which may be suboptimal and consequently limit the quality of the learned representations. We therefore investigate a framework for jointly discovering object and attribute representations. Our key contribution is leveraging the Linear Representation Hypothesis (LRH), which postulates that composable concepts can be represented as linearly additive subspaces in slot representations. Based on this insight, we propose a probabilistic model connecting images, slots (objects), and blocks (attributes). We present an architecture that leverages block attention to connect attribute representations to slots and incorporates LRH in both object and attribute representation spaces. This architecture effectively optimizes the Evidence Lower Bound (ELBO) of the proposed graphical model. Our experiments demonstrate (i) effective discovery of disentangled object and attribute representations, (ii) empirical evidence for LRH in slot space, and (iii) the ability to perform image editing owing to the disentangled and interpretable nature of the learned representations. Our experiments on multiple datasets demonstrate improvements in DCI scores over state-of-the-art methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑