作为令牌化图像生成全局先验的测试时寄存器
Test-Time Registers as Global Priors for Tokenized Image Generation
浏览论文内容
中文总结 AI 辅助
研究在令牌化图像生成中,基于测试时寄存器特征低频集中度高及与DCT低频能量相关的发现,引入RegToken,通过三步将寄存器结构转换为全局先验令牌,插入生成管道可提升指标、加速优化,证明相关结构可作轻量级全局先验。
中文摘要 AI 辅助
基于注意力的模型常出现注意力汇聚问题,少数令牌反复吸引注意并积累大量激活。视觉Transformer中,这些异常值与寄存器密切相关,现有工作多通过可解释性分析和线性探针研究寄存器,未探讨其能否作为即插即用信号用于生成且无需重新训练。本文在令牌化图像生成中重新审视此问题,发现测试时寄存器特征低频集中度更高,与像素空间DCT低频能量有相关性。基于此,引入RegToken,通过三步将寄存器结构转换为全局先验令牌,插入冻结的一维令牌生成管道可提升生成和对齐指标,加速测试时优化。结果表明可将常被视为注意力伪像的结构用作令牌化生成的轻量级全局先验。
英文摘要
Attention-based models often develop attention sinks, where a small number of tokens repeatedly attract attention and accumulate unusually large activations. In vision transformers, these outliers are closely related to registers, which have been diagnostically linked to global, low-frequency image structure. Existing work has largely studied registers through interpretability analyses and linear probes, leaving open whether they can be operationalized as plug-and-play signals for generation without retraining. We revisit this question in tokenized image generation. Using OpenCLIP and DINOv2 on ImageNet, we find that test-time register features exhibit stronger low-frequency concentration than both [CLS] readouts and patch-mean features, and show a consistent (albeit moderate) correlation with pixel-space DCT low-frequency energy. Motivated by these diagnostics, we introduce RegToken, a training-free procedure that converts register structure into a small set of global prior tokens by (i) NFN-based layer localization, (ii) TokenRank-guided subspace extraction, and (iii) a projection-and-conservation update on the register subspace. Inserted into a frozen compact 1D token generation pipeline, RegToken improves ImageNet generation and alignment metrics (e.g., FID-5k 20.5 to 20.1, SigLIP 3.6 to 3.9) without modifying pretrained weights, and accelerates test-time optimization (Steps@$τ$ 74 to 52). Overall, our results suggest that structures often viewed as attention artifacts can be repurposed as lightweight global priors for tokenized generation.
发表机构
- Stony Brook University(纽约州立大学石溪分校)
- Brookhaven National Laboratory(布鲁克海文国家实验室)
机构由 AI 辅助整理,请以论文原文为准。