arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SynVAR:在视觉自回归模型中协同空间与语义对齐

SynVAR: Synergizing Spatial and Semantic Alignment in Visual Autoregressive Model

Zhennan Chen, Tianxing Shi, Pengcheng Xu, Kepan Nan, Qian Wang, Zili Yi, Jian Yang, Ying Tai

arXiv 2608.07948首次发表:更新:

发表机构

State Key Laboratory for Novel Software Technology, Nanjing University; Western University; JIUTIAN Research, CMCC, China(南京大学计算机软件新技术国家重点实验室; 西安大略大学; 中国移动通信集团九天研究)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出首个专为视觉自回归模型(VAR)定制的无训练增强框架SynVAR,通过空间-语义协同控制策略及三个关键组件抑制误差,显著提升了VAR的复杂场景建模能力。

AI 中文摘要

自回归模型(VAR)因其下一个尺度预测范式而获得广泛普及,但在处理包含多个对象和属性的复杂场景时面临严重的性能瓶颈。现有基于扩散的增强方法无法充分解决VAR中跨尺度误差传播与积累的独特挑战。为此,我们提出SynVAR,这是首个专为VAR范式定制的无训练增强框架,引入空间-语义协同控制策略以有效抑制传播误差并提升生成质量。SynVAR包含三个关键组件:(1)全局引导,确保早期阶段合理的空间结构;(2)感受野约束,缓解早期语义混淆;(3)高频补偿,恢复细粒度细节。大量定量与定性实验表明,SynVAR在增强VAR的复杂场景建模能力方面取得显著提升。

英文摘要

VAR has gained widespread popularity due to its next-scale prediction paradigm. However, it faces substantial performance bottlenecks when handling complex scenes with multiple objects and attributes. Existing diffusion-based enhancement methods fail to adequately address the unique challenge of cross-scale error propagation and accumulation in VAR. To this end, we propose SynVAR, the first training-free enhancement framework specifically tailored for the VAR paradigm, which introduces a spatial-semantic collaborative control strategy to effectively suppress propagation error and improve generation quality. SynVAR comprises three key components: (1) Global guidance to ensure reasonable spatial structure in the early stages, (2) Receptive field constraints to mitigate early-stage semantic confusion, (3) High-frequency compensation to recover fine-grained details. Extensive quantitative and qualitative experiments demonstrate the significant improvements in the ability of SynVAR to enhance the VAR's capability for complex scene modeling.

CommentsAccepted by ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑