发表机构
State Key Laboratory for Novel Software Technology, Nanjing University; Western University; JIUTIAN Research, CMCC, China(南京大学计算机软件新技术国家重点实验室; 西安大略大学; 中国移动通信集团九天研究)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出首个专为视觉自回归模型(VAR)定制的无训练增强框架SynVAR,通过空间-语义协同控制策略及三个关键组件抑制误差,显著提升了VAR的复杂场景建模能力。
AI 中文摘要
自回归模型(VAR)因其下一个尺度预测范式而获得广泛普及,但在处理包含多个对象和属性的复杂场景时面临严重的性能瓶颈。现有基于扩散的增强方法无法充分解决VAR中跨尺度误差传播与积累的独特挑战。为此,我们提出SynVAR,这是首个专为VAR范式定制的无训练增强框架,引入空间-语义协同控制策略以有效抑制传播误差并提升生成质量。SynVAR包含三个关键组件:(1)全局引导,确保早期阶段合理的空间结构;(2)感受野约束,缓解早期语义混淆;(3)高频补偿,恢复细粒度细节。大量定量与定性实验表明,SynVAR在增强VAR的复杂场景建模能力方面取得显著提升。
英文摘要
VAR has gained widespread popularity due to its next-scale prediction paradigm. However, it faces substantial performance bottlenecks when handling complex scenes with multiple objects and attributes. Existing diffusion-based enhancement methods fail to adequately address the unique challenge of cross-scale error propagation and accumulation in VAR. To this end, we propose SynVAR, the first training-free enhancement framework specifically tailored for the VAR paradigm, which introduces a spatial-semantic collaborative control strategy to effectively suppress propagation error and improve generation quality. SynVAR comprises three key components: (1) Global guidance to ensure reasonable spatial structure in the early stages, (2) Receptive field constraints to mitigate early-stage semantic confusion, (3) High-frequency compensation to recover fine-grained details. Extensive quantitative and qualitative experiments demonstrate the significant improvements in the ability of SynVAR to enhance the VAR's capability for complex scene modeling.
CommentsAccepted by ECCV 2026