arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25386cs.CV

带前瞻的高效训练:用于自回归图像生成的多令牌辅助监督

Efficient Training with Foresight: Multi-Token Auxiliary Supervision for Autoregressive Image Generation

Guo Niu, Xiongfei Yao, Teng Wang, Nannan Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对自回归图像生成的缺陷,提出MTAR框架,通过多令牌预测、令牌级对比正则化和语义丢弃提升性能,在ImageNet上实现了生成质量与效率的更优平衡。

中文摘要 AI 辅助

自回归(AR)图像生成通过将图像建模为离散令牌序列,在可扩展高保真合成方面展现出强大潜力。然而,传统的下一个令牌预测(NTP)仍存在监督稀疏且短视、表示判别力不足,以及对完整令牌序列进行密集计算导致训练成本高的问题。为解决这些问题,我们提出多令牌自回归(MTAR),这是一个统一的训练框架,从预测目标、表示正则化和训练效率三个方面改进自回归图像生成。具体而言,MTAR引入多令牌预测(MTP),通过对多个未来令牌施加联合监督来缓解传统NTP的稀疏性和短视性;采用令牌级对比正则化(TCR),显式增强采样令牌表示的可分性,从而提高表示判别力;并引入语义丢弃(SD)作为语义感知的训练加速策略,减少低信息令牌的冗余计算,同时保留有价值的学习信号。这三个组件仅在训练期间使用,在自回归推理时不会引入额外开销。在ImageNet上,MTAR在生成质量和训练效率之间实现了更好的平衡。与LlamaGen相比,MTAR的FID低至0.95,训练速度快39%;此外,即使仅用1/3的训练迭代,它仍能达到与基线相当或更优的性能,大幅缩短了训练时间。

英文摘要

Autoregressive (AR) image generation has shown strong potential for scalable high-fidelity synthesis by modeling images as discrete token sequences. However, traditional next token prediction (NTP) continues to suffer from sparse and myopic supervision, insufficiently discriminative representations, and high training cost caused by dense computation over the full token sequence. To address these issues, we propose multi-token autoregressive (MTAR), a unified training framework that improves autoregressive image generation from three aspects: prediction objectives, representation regularization, and training efficiency. Specifically, MTAR introduces multi-token prediction (MTP) to alleviate the sparsity and myopia of traditional NTP by imposing joint supervision on multiple future tokens; employs token-level contrastive regularization (TCR) to explicitly enhance the separability of sampled token representations and thereby improve representation discriminability; and incorporates semantic dropping (SD) as a semantics-aware training acceleration strategy to reduce redundant computation on low-information tokens while preserving informative learning signals. All three components are applied only during training and introduce no additional overhead during autoregressive inference. On ImageNet, MTAR achieves a better balance between generation quality and training efficiency. Compared with LlamaGen, MTAR achieves up to 0.95 lower FID and 39\% faster training. Moreover, even with only 1/3 of the training iterations, it still attains performance comparable to or better than the baseline, substantially reducing training time.

发表机构

  • Foshan University(佛山大学)
  • The University of Hong Kong(香港大学)
  • Sun Yat-sen University(中山大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑