发表机构
University of Wisconsin–Madison; GE HealthCare(威斯康星大学麦迪逊分校; 通用电气医疗)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出DenseAR范式,将自回归图像生成重构成下密集步长预测,解决现有模型局限。基于此扩展为统一模型处理多模态和成像任务,在医学和自然图像验证,在多对比脑MRI及ImageNet上取得良好效果。
AI 中文摘要
我们介绍了DenseAR,一种新的生成范式,它使用紧凑的单尺度分词器将自回归图像生成重新表述为从粗到细的下密集步长预测。关键在于以逐渐密集的步长遍历单尺度潜在网格能自然捕捉从全局结构到精细细节的过渡。这同时解决了现有自回归模型的两个局限:光栅顺序自回归推理慢,DenseAR通过并行预测多个令牌避免;多尺度方法成本高,因其需长的多分辨率令牌序列来实现从粗到细预测。基于高效框架和自回归建模灵活性,我们将DenseAR扩展为统一模型以处理多种模态和成像任务。我们在医学和自然图像上验证了DenseAR。在多对比脑MRI上,单个DenseAR模型统一了跨模态翻译、模态条件生成和肿瘤分割,且与特定任务方法竞争。在ImageNet上,DenseAR在类条件生成质量(FID和IS)上优于无步长排序的单网格基线和基于多尺度分词器的基线。
英文摘要
We introduce DenseAR, a new generative paradigm that reformulates autoregressive image generation as coarse-to-fine next-dense-stride prediction using a compact single-scale tokenizer. Our key insight is that traversing a single-scale latent grid with progressively denser strides naturally captures the transition from global structure to fine detail. This addresses two limitations of existing autoregressive models at once: the slow inference of raster-order autoregression, which DenseAR avoids by predicting multiple tokens in parallel, and the heavy cost of multi-scale approaches, which need long, multi-resolution token sequences to achieve coarse-to-fine prediction. Building on our efficient framework and the flexibility of autoregressive modeling, we further extend DenseAR to a unified model that handles multiple modalities and imaging tasks within a single backbone. We validate DenseAR on both medical and natural images. On multi-contrast brain MRI, a single DenseAR model unifies cross-modal translation, modality-conditioned generation, and tumor segmentation, while remaining competitive with task-specific methods. On ImageNet, DenseAR improves class-conditional generation quality (FID and IS) over both a single-grid baseline without stride ordering and a multi-scale tokenizer-based baseline.