arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LoopVAE:跨尺度的循环深度用于视觉分词

LoopVAE: Recurrent Depth Across Scales for Visual Tokenization

Zhiying Lu

arXiv 2609.11516首次发表:更新:

发表机构

University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

LoopVAE通过跨尺度重用循环核心,在减少65%参数的同时,在ImageNet-256上达到0.28 rFID和32.54 dB PSNR,确立了循环深度作为视觉分词中参数共享的设计轴。

AI 中文摘要

层次化视觉分词器通常将不同的处理块分配给不同的空间尺度。我们探究了这种计算中有多少可以使用相同的参数。LoopVAE在尺度内部和跨尺度之间重用了一个尺度与循环条件化的核心,同时保持分辨率变化的转换独立。一个四块核心在每次编码器或解码器中执行28次块应用。在ImageNet-256上,这个29M参数的卷积模型在约30个epoch的两阶段训练预算下达到了0.28 rFID和32.54 dB PSNR,比84M参考VAE少用了约65%的参数。一个具有相同执行图的非对抗性Transformer消融研究发现,在全局共享下PSNR和SSIM具有竞争力,尽管非共享块改善了LPIPS。针对性的循环干预表明,完成训练好的循环可以改善重建,即使小的特征更新也可能产生显著的下游影响。截断还暴露了输出范围错误,区分了有用的循环计算与可靠的提前退出。运行时分析揭示了执行权衡:在测试配置中,存储更少的权重需要更多的算术运算和更长的运行时间。通过卷积和Transformer算子以及单分辨率或多分辨率潜接口,LoopVAE确立了跨尺度的循环深度作为视觉分词的一个参数共享设计轴。

英文摘要

Hierarchical visual tokenizers typically allocate different processing blocks to different spatial scales. We ask how much of this computation can use the same parameters. LoopVAE reuses a scale- and loop-conditioned core within and across scales, while keeping resolution-changing transitions independent. A four-block core executes 28 block applications per encoder or decoder. On ImageNet-256, the 29M-parameter convolutional model reaches 0.28 rFID and 32.54 dB PSNR under an approximately 30-epoch two-stage training budget, using approximately 65% fewer parameters than the 84M reference VAEs. A non-adversarial Transformer ablation with the same execution graph finds competitive PSNR and SSIM under global sharing, although unshared blocks improve LPIPS. Targeted loop interventions show that completing the trained recurrence improves reconstruction and that even small feature updates can have substantial downstream effects. Truncation also exposes output-range errors, distinguishing useful recurrent computation from reliable early exit. Runtime profiling reveals the execution tradeoff: fewer stored weights require more arithmetic and longer runtime in the tested configurations. With convolutional and Transformer operators and single- or multi-resolution latent interfaces, LoopVAE establishes recurrent depth across scales as a parameter-sharing design axis for visual tokenization.

Comments19 pages, 6 figures, 7 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑