arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11804cs.CVcs.AIcs.LG

Logit Refiner:通过尺度内依赖建模改进视觉自回归模型

Logit Refiner: Improving Visual Autoregressive Models via Intra-Scale Dependency Modeling

Meimingwei Li, Stefan Andreas Baumann, Felix Krause, Björn Ommer

首次发表
浏览论文内容

中文总结 AI 辅助

针对视觉自回归模型并行解码忽略尺度内依赖导致局部不连贯的问题,提出Logit Refiner轻量模块,通过顺序采样恢复依赖,仅增10%参数,显著提升生成质量并泛化至文本到图像生成。

中文摘要 AI 辅助

视觉自回归模型(VAR)通过下一尺度预测生成图像,并行生成每个尺度内的所有标记。我们表明,这种并行解码构成了一种平均场风格的近似,丢弃了同尺度标记之间的空间依赖关系,导致无论骨干网络容量如何,样本都缺乏局部连贯性——这是解码规则本身的局限性。为解决这一局限性,我们引入了Logit Refiner,一个轻量级自回归模块,通过基于冻结的骨干特征顺序采样标记来恢复尺度内依赖关系。仅增加约10%的参数和少于基础模型5%的训练计算量,它即可插入任何预训练的VAR检查点而无需重新训练。受控消融实验表明,联合尺度内采样——而非额外容量或训练——是关键因素。在类别条件ImageNet 256x256上,从310M到2B参数的骨干网络中,该精炼器持续提升生成质量,使1.1B参数的模型超越两倍大小的模型。该方法进一步推广到文本到图像生成,证实平均场瓶颈在VAR变体中持续存在,并被我们的方法有效缓解。项目页面:此https URL

英文摘要

Visual Autoregressive Models (VAR) generate images through next-scale prediction, producing all tokens within each scale in parallel. We show that this parallel decoding constitutes a mean-field-style approximation that discards spatial dependencies among same-scale tokens, causing locally incoherent samples regardless of backbone capacity -- a limitation of the decoding rule. Addressing this limitation, we introduce the Logit Refiner, a lightweight autoregressive module that restores intra-scale dependencies by sequentially sampling tokens conditioned on frozen backbone features. Adding only ~10% parameters and less than 5% of the base model's training compute, it plugs into any pretrained VAR checkpoint without retraining. Controlled ablations isolate joint intra-scale sampling -- rather than additional capacity or training -- as the critical ingredient. Across backbones from 310M to 2B parameters on class-conditional ImageNet 256x256, the refiner consistently improves generation quality, enabling a 1.1B-parameter model to surpass one twice its size. The approach further generalizes to text-to-image generation, confirming that the mean-field bottleneck persists across VAR variants and is effectively alleviated by our method. Project page: https://compvis.github.io/logit-refiner/

发表机构

  • Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑