arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Slot-RAE:通过直接表示自动编码器简化以对象为中心的学习

Slot-RAE: Streamlining Object-Centric Learning via Direct Representation Auto-Encoders

Alexandre Chapin, Emmanuel Dellandrea, Liming Chen

arXiv 2607.11196首次发表:更新:

发表机构

LIRIS; Ecole Centrale de Lyon(里昂图像与信息系统实验室; 里昂中央理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对以对象为中心的模型部署问题,提出Slot-RAE框架,直接在视觉基础模型特征空间运行,采用特征空间扩散过程及独特训练方式,在COCO数据集实验中取得领先成果,实现高效无监督对象发现等,速度和计算效率更高。

AI 中文摘要

为现实世界场景理解部署以对象为中心的模型通常需要复杂的管道来实现强大的场景分解和高保真生成。最近基于扩散的方法提高了视觉质量,但几乎普遍依赖繁重的预训练生成先验(如Stable Diffusion)和外部VAE潜在空间。本文提出Slot-RAE,一个更简单、完全集成的框架,直接在视觉基础模型(如DINOv3)的连续语义特征空间内运行。它采用基于特征空间的扩散过程,使用扩散变压器解码器和表示对齐头。与现有基于扩散的以对象为中心的方法不同,Slot-RAE的生成核心在冻结的VFM特征空间内从头开始训练,无需VAE瓶颈和与任务无关的生成预训练。在COCO数据集上的实验表明,Slot-RAE虽架构简单,但取得了领先成果,在无监督对象发现、图像重建和零样本合成性方面表现出色,且速度更快、计算效率更高。

英文摘要

Deploying object-centric models for real-world scene understanding typically requires complex pipelines to achieve both robust scene decomposition and high-fidelity generation. Recent diffusion-based approaches have improved visual quality, but they almost universally rely on heavy, pretrained generative priors (e.g., Stable Diffusion) and external VAE latent spaces. In this paper, we propose Slot-RAE, a much simpler, fully integrated framework that operates directly within the continuous semantic feature space of visual foundation models (e.g., DINOv3). Slot-RAE employs a feature-space diffusion process using a Diffusion Transformer (DiT) decoder and a Representation Alignment (REPA) head. Unlike existing diffusion-based objectcentric methods that rely heavily on subsidized text-toimage priors, the generative core of Slot-RAE (Slot Attention and the DiT) is trained from scratch within the frozen VFM feature space. This eliminates the need for VAE bottlenecks and task-agnostic generative pre-training. Experiments on the COCO dataset demonstrate that despite its architectural simplicity, Slot-RAE achieves state-of-the-art results. It delivers comparable unsupervised object discovery, higher-fidelity image reconstruction, and robust zero-shot compositionality, all while being significantly faster and more computationally efficient than existing object-centric latent diffusion models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑