VISTA:面向视觉自回归生成的测试时组合对齐方法
VISTA: Test-Time Compositional Alignment for Visual Autoregressive Generation
浏览论文内容
中文总结 AI 辅助
VISTA是首个面向视觉自回归图像生成的测试时对齐框架,基于Infinity构建,可提升组合生成能力,在2B、8B模型上分别将平均目标分数提近20%、近6%,且不损失图像质量。
中文摘要 AI 辅助
视觉自回归(VAR)模型已成为文本到图像生成中扩散模型的快速、高质量替代方案,但与扩散模型类似,它们仍存在持续的组合失败问题,生成的图像会违反提示中指定的属性绑定和空间关系。尽管针对扩散模型已开发了大量测试时对齐方法,但针对下一代VAR生成的可比方法仍不存在,VAR的有状态、离散、多分辨率采样过程使得现有技术无法适用。我们通过VISTA(Visual Autoregressive Semantic Test-time Alignment,视觉自回归语义测试时对齐)填补了这一空白,它是首个面向下一代自回归图像生成的基于梯度的测试时对齐框架。VISTA基于Infinity构建,直接干预生成过程,通过冻结的Transformer优化中间表示,以引导视觉预测符合组合约束,无需修改模型参数或额外训练。VISTA引入了使此类优化在不同尺度上保持稳定所需的机制,以及一个可扩展的目标空间,任何对交叉注意力的可微约束都可接入该空间。在两个基准和两个模型规模上,VISTA提升了所有目标组合类别,在2B主干模型上将平均目标分数提高了近20%,在8B主干模型上提高了近6%,在空间关系上获得最大增益。图像质量得以保留:一个VISTA从未优化的独立偏好模型对其输出的评分高出近20%。值得注意的是,配备VISTA的2B模型超过了其规模四倍的主干模型,表明模型尺度间的组合差距有相当一部分可在测试时恢复。
英文摘要
Visual autoregressive (VAR) models have emerged as a fast, high-quality alternative to diffusion for text-to-image generation, but like diffusion models they exhibit persistent compositional failures, producing images that violate the attribute bindings and spatial relations specified in the prompt. While a rich line of test-time alignment methods has developed for diffusion, no comparable approach exists for next-scale VAR generation, whose stateful, discrete, multi-resolution sampling process makes existing techniques inapplicable. We close this gap with \textbf{VISTA} (\textbf{Vi}sual Autoregressive \textbf{S}emantic \textbf{T}est-time \textbf{A}lignment), the first gradient-based test-time alignment framework for next-scale autoregressive image generation. Built on Infinity, VISTA intervenes directly in the generation process, optimizing intermediate representations through the frozen transformer to steer visual predictions toward compositional constraints, without modifying model parameters or requiring additional training. VISTA introduces the mechanisms needed to make such optimization stable across scales, together with an extensible objective space that any differentiable constraint on cross-attention can plug into. Across two benchmarks and two model scales, VISTA improves every targeted compositional category, raising the mean targeted score by nearly 20\% on a 2B backbone and almost 6\% on an 8B backbone, with the largest gains on spatial relations. Image quality is preserved: an independent preference model VISTA never optimizes scores its outputs nearly 20\% higher. Notably, the 2B model with VISTA surpasses a backbone four times its size, indicating that a substantial part of the compositional gap between model scales is recoverable at test time.
发表机构
- Sharif University of Technology(谢里夫理工大学)
机构由 AI 辅助整理,请以论文原文为准。