arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

程序化核心:一种用于视觉Transformer的紧凑循环初始化方法

Procedural Core: A Compact Recurrent Initialization for Vision Transformers

Zachary Shinnick, Christian Internò, Hemanth Saratchandran, Anton van den Hengel, Damien Teney

arXiv 2609.37631首次发表:更新:

发表机构

Australian Institute for Machine Learning (AIML), Adelaide University; Bielefeld University; Metacognition AI; Idiap Research Institute(阿德莱德大学澳大利亚机器学习研究所; 比勒费尔德大学; Metacognition AI; Idiap 研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出程序化核心初始化方法,用极小循环Transformer在程序化数据上训练紧凑权重,扩展后可初始化任意ViT,在图像分类、自监督学习等任务上超越随机初始化,证明Transformer无需从零开始。

AI 中文摘要

Transformer通常从随机初始化开始训练,这要求其所有能力都从大规模优化中涌现。近期研究表明,少量抽象的程序化生成数据有助于以低成本获取通用的归纳结构。然而,这增加了一个必须为每个目标模型重复的预训练阶段。我们提出程序化核心(Procedural Core),一种初始化策略,将这种通用结构捕获到一组紧凑的权重中,这些权重可以在不同模型间复用。我们在程序化数据上训练一个极小的循环Transformer,然后扩展其权重以初始化任意宽度和深度的Transformer。由此产生的初始化在图像分类、自监督视觉学习(DINO)以及自然语言建模(FineWeb-Edu)和代码建模(CodeParrot)上均提升了性能。对于图像分类,将一个1M参数的核心扩展以初始化85M参数的ViT-Base,在ImageNet top-1准确率上比标准随机初始化提高了2.2个百分点。我们的分析表明,循环对于学习可跨模型迁移的紧凑权重至关重要。在ViT中,我们将一个关键优势定位于抑制高范数token,这带来了零样本分割(ImageNet-S mAP从32.3提升至42.9)、目标定位(VOC07 CorLoc从9.9提升至18.4)和深度估计(NYUv2 RMSE从1.104降至0.998)的显著改进。这表明Transformer不必从空白状态开始,可以以低成本、无需任何领域或任务特定数据的方式,用通用能力进行初始化。

英文摘要

Transformers are typically trained from random initialization, requiring all their capabilities to emerge from large-scale optimization. Recent work showed that a small amount of abstract procedurally generated data can help acquire generic inductive structure at low cost. However, this adds a pretraining stage that must be repeated for every target model. We propose Procedural Core, an initialization strategy that captures this generic structure into a compact set of weights that can be reused across models. We train a minimal recurrent transformer on procedural data, then expand its weights to initialize transformers of arbitrary width and depth. The resulting initialization improves performance on image classification, self-supervised visual learning (DINO), and modeling natural language (FineWeb-Edu) and code (CodeParrot). For image classification, expanding a 1M-parameter core to initialize an 85M-parameter ViT-Base improves ImageNet top-1 accuracy by 2.2 pp over standard random initialization. Our analysis identifies recurrence as essential for learning compact weights that transfer across models. In ViTs, we localize a key benefit in the suppression of high-norm tokens that produces substantial improvements in zero-shot segmentation (ImageNet-S mAP 32.3 to 42.9), object localization (VOC07 CorLoc 9.9 to 18.4), and depth estimation (NYUv2 RMSE 1.104 to 0.998). This demonstrates that transformers need not start from a blank slate, and can be initialized with generic capabilities at low cost with no domain- or task-specific data.

CommentsProject page: zlshinnick.github.io/procedural-core/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑