arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39657cs.CV

结构无关蒸馏的可扩散潜变量

Diffusable Latents from Structure-Agnostic Distillation

Adrien Ramanana Rahary, Nicolas Dufour, Patrick Pérez, David Picard

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出结构无关蒸馏方法,通过池化图像级描述符对齐替代密集位置级对齐,提升潜变量可扩散性,并验证一阶与关系目标在多种模态下的有效性。

中文摘要 AI 辅助

将预训练基础模型蒸馏到自编码器瓶颈中可提高潜变量的可扩散性,使扩散模型更快收敛并达到更高的样本质量。标准蒸馏将每个位置的潜变量与同位置的教师特征对齐,从而将潜变量布局与教师布局绑定。我们证明这一约束并非必要:将单个池化的图像级描述符与教师对齐,其性能与密集的位置级蒸馏相当或略优。我们比较了不同潜变量形状和教师模态下的一阶目标与关系目标。一阶匹配可自然扩展到一维令牌序列潜变量及跨模态场景,其中将文本编码器蒸馏到图像自编码器中仍能提高可扩散性;仅基于每幅图像最近邻的关系目标也能提升可扩散性。代码和博客文章可在以下网址获取:此 https 链接和此 https 链接。

英文摘要

Distilling pretrained foundation models into an autoencoder bottleneck improves latent diffusability, enabling diffusion models to converge faster and reach higher sample quality. Standard distillation aligns the latent at each position to a co-located teacher feature, tying the latent layout to the teacher's. We show this constraint is unnecessary: aligning a single pooled image-level descriptor to the teacher's performs as well as or slightly better than dense position-wise distillation. We compare first-order and relational pooled objectives across latent shapes and teacher modalities. First-order matching extends naturally to 1D token-sequence latents and across modalities, where distilling a text encoder into an image autoencoder still improves diffusability; a relational objective based only on each image's nearest neighbours improves it as well. Code and blog post are available at https://github.com/AdrienRR/structure-agnostic-distillation and https://kyutai.org/blog/2026-09-28-structure-agnostic-distillation/.

发表机构

  • LIGM
  • ENPC(巴黎高科路桥学校)
  • IP Paris(巴黎综合理工学院)
  • CNRS(法国国家科学研究中心)
  • UGE(巴黎东大学)
  • Kyutai

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑