发表机构
University of Würzburg; Google(维尔茨堡大学; 谷歌)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对像素空间扩散Transformer的令牌冗余问题,提出Region Token Interface弹性令牌压缩方法,可在多压缩预算下提升效率,2.0倍速时达密集模型质量,2.6倍速时仍接近该质量。
AI 中文摘要
自然图像的细节集中在画面的一小部分区域,但扩散模型在每一层、每一时间步都会为每个图像补丁分配一个完整的令牌,这种资源浪费在像素空间模型中最为严重,因为这类模型没有自动编码器来预先吸收低级冗余。本文通过探究一个预训练的像素空间文本到图像Transformer,发现其中间块的令牌在图像平坦区域是冗余的,这种冗余占据了与内容形状一致的连通区域,要利用该冗余需要具有相同几何结构的令牌,而对补丁进行希尔伯特排序的切割正好能提供这类令牌:连续位置始终是图像的相邻位置,因此任何连续序列都是大小和形状随内容变化的连通区域,二维分组可转化为一维切割。现有压缩方法各有缺陷:相似性合并会分散分组,潜在瓶颈会丢弃位置信息,跳步会删除本应汇总的内容。本文在模型特征变化最剧烈的位置进行切割,并将每个序列池化为一个区域令牌,提出了区域令牌接口(Region Token Interface,即\method{}),该接口可使扩散模型适配这类令牌,且在微调期间随机抽取区域数量,因此一个检查点可适配所有压缩预算。实验表明,在相同压缩预算下,\method{}的性能优于现有压缩方法;在2.0倍速度下可达到密集模型的质量,在2.6倍速度下仍接近密集模型的质量。代码和模型已在指定URL开源。
英文摘要
Natural images concentrate their detail in a small fraction of the frame, yet diffusion models spend a full token on every patch, in every layer and at every timestep. The waste is largest in pixel-space models, with no autoencoder to absorb low-level redundancy first. Probing a pretrained pixel text-to-image transformer, we find its middle-block tokens redundant wherever the image is flat. The redundancy occupies connected, content-shaped regions, and exploiting it requires tokens with the same geometry. Cutting a Hilbert ordering of the patches provides them. Consecutive positions are always image neighbours, so any contiguous run is a connected region whose size and shape follow the content, and grouping in two dimensions becomes a cut in one. Existing reductions each lose part of this. Similarity merging scatters its groups, latent bottlenecks discard position, and skipping deletes what it should summarize. We cut where the model's features change most and pool each run into one region token. Our Region Token Interface (\method{}) adapts a diffusion model to these tokens, with the region count drawn at random during fine-tuning so one checkpoint serves every budget. \method{} leads prior reduction methods at matched budgets, matches dense quality at $2.0\times$ the speed, and stays close at $2.6\times$. The code and models are open-sourced at https://eduardzamfir.github.io/rti