发表机构
University of Illinois Urbana-Champaign; Google(伊利诺伊大学厄巴纳-香槟分校; 谷歌公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
BudgetPix是一种自适应分词框架,通过动态分配算力优化图像扩散模型的质量-效率权衡,仅用25%原预算即可达到对应模型的保真度,性能优于现有基线。
AI 中文摘要
大多数图像生成模型依赖均匀分词,为等大小的图像块分配完全相同的计算预算。这种静态范式无法在推理时适应不同的资源约束,且会在纯色背景和复杂细节上投入相同的算力,导致质量-成本权衡不佳。我们提出BudgetPix,这是一种自适应分词框架,可基于视觉复杂度和空间布局动态分配算力,实现推理时的灵活计算预算控制。BudgetPix包含三个关键组件:(1)自适应编码器,使用熵引导的四叉树结合多尺度块嵌入器,将固定大小的图像映射为可变长度的分词序列;(2)尺度感知解码器,从多尺度分词集中重建固定分辨率图像;(3)灵活的训练与采样调度,使像素空间去噪器能在可变分词数量下运行。BudgetPix可无缝集成到现有像素空间扩散架构中,使单个检查点能在广泛的计算预算下运行。在文本到图像生成任务中,BudgetPix仅使用原始计算预算的25%,就达到了MiniT2I-L在512²分辨率和PixelDiT在1024²分辨率下的保真度。在使用MeanFlow骨干网络的类别条件生成中,BudgetPix仅需60%的全量计算预算,就能生成质量几乎无下降的图像,FID仅上升0.8个点。人类和视觉语言模型(VLM)的综合评估证实,与现有的预算自适应基线相比,BudgetPix实现了显著更优的质量-效率权衡。更多详情可访问我们的项目页面:this https URL
英文摘要
Most image generation models rely on uniform tokenization, allocating the exact same computational budget to equally-sized image patches. This static paradigm cannot adapt to different resource constraints at inference time, and yields suboptimal quality-cost tradeoff by devoting the same effort to both plain backgrounds and intricate details. We propose BudgetPix, an adaptive tokenization framework that dynamically allocates compute based on visual complexity and spatial layout, enabling flexible computational budgeting at inference time. BudgetPix comprises three key components: (1) an adaptive encoder that maps a fixed-size image to a variable-length token sequence using an entropy-guided quadtree alongside a multi-scale patch embedder; (2) a scale-aware decoder reconstructs fixed-resolution images from multi-scale token sets; and (3) a flexible training and sampling schedule that enables pixel-space denoisers to operate across variable token counts. BudgetPix seamlessly integrates with existing pixel-space diffusion architectures, enabling a single checkpoint to be operated at a wide range of compute budgets. Evaluated on text-to-image generation, BudgetPix matches the fidelity of MiniT2I-L at $512^2$ and PixelDiT at $1024^2$ using just 25% of the original compute budget. In class-conditional generation using a MeanFlow backbone, BudgetPix requires merely 60% of the full compute budget to produce images with near-zero quality degradation, observing a marginal 0.8-point increase in FID. Comprehensive assessments by human and VLM judges confirm that BudgetPix establishes a significantly improved quality-efficiency tradeoff over prior budget-adaptive baselines. More details are available at our project page: https://karaozgur.com/BudgetPix
CommentsMore details are available at our project page: https://karaozgur.com/BudgetPix