arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Token扩散的层级(Level-of-Token Diffusion)

Level-of-Token Diffusion

Kiyohiro Nakayama, Brian Chao, Jan Ackermann, Hansheng Chen, Federico Tombari, Leonidas Guibas, Lior Yariv, Gordon Wetzstein

arXiv 2610.05816首次发表:更新:

发表机构

Stanford University; Google(斯坦福大学; 谷歌)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Token扩散的层级(LoT Diffusion),通过显式多分辨率Token布局实现自适应计算分配,在保留预训练生成先验的同时,显著提升图像和视频生成的效率与质量平衡。

AI 中文摘要

图像和视频扩散模型对每个区域分配相同的计算量,即使目标场景需要不同级别的细节。细节的空间分布通常可以在生成之前预见到,从而指示出哪些地方可以减少计算。我们引入了Token扩散的层级(LoT Diffusion),这是一个框架,将这种知识转化为显式的多分辨率Token布局(Level-of-Token布局),用于自适应和高效的生成。Token代表不同大小和形状的矩形补丁,在需要细节的地方分配更精细的Token,在其他地方分配更粗糙的Token。我们通过补丁级非对称流参数化和多分辨率Token的嵌入,将预训练的扩散Transformer适配到LoT布局,在每一步去噪过程中保持全分辨率流预测,同时仅处理减少的Token序列。LoT Diffusion实现了布局自适应生成,同时保留了预训练的生成先验。我们使用从语义掩码、边界框、纹理方差和景深线索以及智能体计划中得出的布局来演示LoT。在图像和视频生成中,LoT提供了有利的质量-效率权衡,显著的加速取决于布局的Token预算。我们的项目网站位于此https URL。

英文摘要

Image and video diffusion models allocate equal computation to every region, even when the intended scene calls for varying levels of detail. The spatial distribution of detail can often be anticipated before generation, indicating where computation can be reduced. We introduce Level-of-Token (LoT) Diffusion, a framework that turns this knowledge into an explicit multiresolution token layout (Level-of-Token layout) for adaptive and efficient generation. Tokens represent rectangular patches of varying sizes and shapes, allocating finer tokens where detail is needed and coarser tokens elsewhere. We adapt pretrained diffusion transformers to LoT layouts through a patch-wise asymmetric flow parametrization and embeddings for multiresolution tokens, preserving full-resolution flow prediction at every denoising step while processing only a reduced token sequence. LoT Diffusion enables layout-adaptive generation while preserving pretrained generative priors. We demonstrate LoT with layouts derived from semantic masks, bounding boxes, texture variance, and depth-of-field cues, as well as agentic plans. Across image and video generation, LoT offers favorable quality-efficiency tradeoffs, with significant speedups determined by the layout's token budget. Our project website is at https://georgenakayama.github.io/lotdiffusion/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑