MOSAIK:用于高效生成的图像令牌多补丁内容感知空间分配
MOSAIK: Multi-Patch Content-Aware Spatial Allocation of Image Tokens for Efficient Generation
浏览论文内容
中文总结 AI 辅助
MOSAIK是一种损伤引导的异构补丁布局框架,适配PixelDiT骨干,通过区域分配补丁大小,在大幅降低计算量和令牌数的同时,保持与全计算PixelDiT相当的生成性能,且在高计算约束下优于其他效率方法。
中文摘要 AI 辅助
像素空间扩散模型通过直接在图像空间生成,避免了潜在扩散模型的重建上限,但由于自注意力的二次复杂度,其显著更高的令牌数导致生成成本高昂。现有几种效率方法通过在选定的去噪步骤使用更大的补丁来减少此成本,从而用更少的令牌表示图像,但每个步骤仍在整个图像上统一使用单一补丁大小,忽略了不同区域在 coarsening(粗化)时会遭受不同的保真度损失。我们引入MOSAIK,一种损伤引导框架,可在区域和去噪步骤间改变补丁大小。MOSAIK适配PixelDiT骨干网络以生成任意异构补丁布局,一个轻量预测器利用中间去噪特征估计每个区域粗化导致的保真度损失。在给定令牌预算下,该损伤引导布局预测器为敏感区域分配精细补丁,其余区域分配粗糙补丁。值得注意的是,MOSAIK在将FLOPs减少70%、令牌数减少83%的同时,在GenEval上与全计算PixelDiT性能匹配,其DPG-Bench评分仅下降1.0点。与时间补丁调度、特征缓存等多种效率范式相比,我们的方法在中等预算下提供极具竞争力的性能,且在高度受限的计算 regime(机制)中始终优于这些基线。
英文摘要
Pixel-space diffusion models avoid the reconstruction ceiling of latent diffusion models by generating directly in image space. However, their substantially higher token count makes generation expensive due to the quadratic complexity of self-attention. Several existing efficiency methods reduce this cost by using larger patches at selected denoising steps, thereby representing the image with fewer tokens. Yet, each step still uses a single patch size uniformly across the entire image, overlooking that different regions suffer different fidelity losses when coarsened. We introduce MOSAIK, a damage-guided framework that varies patch size across regions and denoising steps. MOSAIK adapts the PixelDiT backbone to generate arbitrary heterogeneous patch layouts, and a lightweight predictor uses intermediate denoising features to estimate the fidelity loss caused by coarsening each region. Given a token budget, our damage-guided layout predictor assigns fine patches to sensitive regions and coarse patches elsewhere. Remarkably, while reducing FLOPs by 70% and token count by 83%, MOSAIK matches the full-compute PixelDiT on GenEval and its DPG-Bench score drops by only 1.0 point. Compared to diverse efficiency paradigms, including temporal patch scheduling and feature caching, our approach delivers highly competitive performance at moderate budgets and consistently outperforms these baselines in highly constrained compute regimes.