arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越 token 级交叉熵:用于自回归图像生成的 Fréchet 分布后训练

Beyond Token-Level Cross-Entropy: Fréchet Distributional Post-Training for Autoregressive Image Generation

Jinhua Zhang, Yisong Lin, Wei Long, Shuhang Gu

arXiv 2608.00562首次发表:更新:

发表机构

UESTC(电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对自回归图像生成器预训练与评估的目标及上下文不匹配问题,提出 FD-loss 后训练方法,在 ImageNet 256×256 任务上大幅降低 FID 与 FD_r6,提升生成质量且无额外参数或推理步骤。

AI 中文摘要

自回归图像生成器通常在教师强制机制下以 token 级交叉熵进行预训练,但评估时却以解码图像的分布质量为标准,这造成了目标不匹配——分类错误会产生不等的图像级后果,以及上下文不匹配——推理依赖模型生成的历史序列。我们提出 FD-loss 后训练方法,该方法以表示空间 Fréchet 距离为唯一目标,对预训练的离散生成器进行适配。双通方案首先在模型原生推理配置下通过无梯度生成构建分离的回滚上下文,再通过概率级直通估计器(STE)执行可微重放,该估计器在前向传播中保留硬 argmax 解码,同时通过温度缩放后的概率传播图像级梯度。仅更新生成器,tokenizer 和特征提取器保持冻结状态。在针对类条件 ImageNet 256×256 数据集的四个生成器家族的八个完整配置中,FD-loss 后训练平均将 FID 和 FD_r6 分别降低 41.4% 和 52.0%,最佳 FID 结果从 2.42 提升至 1.43,且未增加参数或推理步骤。

英文摘要

Autoregressive image generators are commonly pretrained with token-level cross-entropy under teacher forcing, yet evaluated by the distributional quality of decoded images. This creates an objective mismatch, because categorical errors have unequal image-level consequences, and a context mismatch, because inference conditions on model-generated histories. We introduce FD-loss post-training, which adapts a pretrained discrete generator using representation-space Fréchet distance as the sole objective. A dual-pass scheme first constructs detached rollout contexts through gradient-free generation under the model's native inference configuration, then performs differentiable replay with a probability-level straight-through estimator (STE) that preserves hard argmax decoding in the forward pass while propagating image-level gradients through temperature-scaled probabilities. Only the generator is updated, while the tokenizer and feature extractors remain frozen. Across eight completed configurations from four generator families on class-conditional ImageNet at $256\times256$, FD-loss post-training reduces FID and $\mathrm{FD}_{r6}$ by 41.4% and 52.0% on average. The strongest FID result improves from 2.42 to 1.43 without adding parameters or inference steps.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑