发表机构
University of Bonn(波恩大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SlotDiT提出以对象为中心的槽作为扩散Transformer的潜在表示,用于文本引导的机器人视频生成,在保持生成质量的同时提升任务完成率并降低计算成本。
AI 中文摘要
文本条件下的潜在扩散模型在视频生成中表现强劲,并且是机器人应用的有前景的骨干网络。然而,现有方法依赖于像素级或基于VAE的潜在表示,这些表示缺乏显式的语义结构,导致表示空间的影响在很大程度上未被探索。基于槽的以对象为中心的表示通过将场景分解为对象级潜在变量(即槽)提供了一种结构化替代方案。虽然这些表示在动力学建模和规划中已显示出成功,但尚未被探索用于基于扩散的生成建模。我们引入了SlotDiT,一种在基于槽的潜在空间中运行的文本引导扩散Transformer(DiT)。给定参考图像和语言指令,SlotDiT将场景分解为代表各个实体的以对象为中心的槽。在指令和观察到的场景上下文的条件下,模型自回归地去噪未来的槽轨迹以预测场景动力学。为了系统地研究扩散Transformer的潜在空间设计,我们在统一的DiT框架内将基于槽的表示与基于VAE和语义对齐的替代方案进行比较。我们的实验表明,使用槽作为DiT潜在变量可产生具有竞争力的视频生成质量,同时在四个机器人数据集上持续提高任务完成率。此外,其紧凑的表示提供了一种计算效率高的替代方案,优于基于VAE和语义对齐的潜在空间。总的来说,我们的结果表明,以对象为中心的结构是机器人环境中基于扩散的生成建模的强大归纳偏置。项目页面可在该https URL获取。
英文摘要
Text-conditioned latent diffusion models perform strongly in video generation and are promising backbones for robotic applications. However, existing approaches rely on pixel-level or VAE-based latent representations that lack explicit semantic structure, leaving the impact of the representation space largely unexplored. Slot-based object-centric representations offer a structured alternative by decomposing scenes into object-level latents, or slots. While they have shown success in dynamics modeling and planning, they have not yet been explored for diffusion-based generative modeling. We introduce SlotDiT, a text-guided Diffusion Transformer (DiT) that operates in a slot-based latent space. Given a reference image and a language instruction, SlotDiT decomposes the scene into object-centric slots representing individual entities. Conditioned on the instruction and observed scene context, the model autoregressively denoises future slot trajectories to predict scene dynamics. To systematically investigate latent-space design for diffusion transformers, we compare slot-based representations against VAE-based and semantics-aligned alternatives within a unified DiT framework. Our experiments show that using slots as DiT latents yields competitive video generation quality while consistently improving task-completion rates across four robotic datasets. Furthermore, their compact representation provides a computationally efficient alternative to VAE-based and semantics-aligned latent spaces. Overall, our results demonstrate that object-centric structure is a powerful inductive bias for diffusion-based generative modeling in robotic environments. The project page is available at https://slot-dit.github.io/.
CommentsAccepted at BMVC 2026. Project page: https://slot-dit.github.io/