AI 中文总结
本文提出SCG方法,使JEPA世界模型的视觉Transformer编码器可按需连续扩展宽度或深度,在多任务上提升性能并实现更高参数效率,且零误扩展。
AI 中文摘要
用于世界建模的联合嵌入预测架构(JEPAs)通常采用固定规模的视觉Transformer编码器,这类编码器对于简单任务而言容量过剩,而对于复杂任务则容量不足,且注意力头之间存在大量冗余。我们提出了连续容量增长(Successive Capacity Growth,SCG)方法,该方法从最小规模编码器(1个注意力头、2层、28.3万参数)开始,在与任务无关的测试验证机制驱动下,根据需求逐步扩展宽度(添加注意力头以提升低级语义容量)或深度(添加Transformer块以实现高阶语义抽象),该机制利用函数保持性扩展来安全地尝试架构变更,若变更未降低预测损失则回滚。绘制各向同性高斯正则化器(Sketched Isotropic Gaussian Regularizer,SIGReg)确保所有学习到的语义维度在统计上保持独立并与预测目标对齐,即使架构扩展也能防止崩溃。在60维多目标动力学任务上,SCG自然触发深度扩展,相比固定小型基准模型,预测损失降低20.3%,且参数效率比扩展至固定大型模型高56倍;在2D导航任务上,单次宽度扩展甚至比固定大型模型性能提升23%。在所有三个测试的复杂度递增环境中,自适应编码器的性能与固定小型基准模型相当或更优,且零误扩展,实现了比特级精确函数保持(比率=1.0,绝对差值=0.0)。核心结论是,JEPA世界模型编码器无需预先分配最大容量,可根据任务需求连续增长,在保持表示质量的同时实现显著的计算与数据效率。
英文摘要
Joint-Embedding Predictive Architectures (JEPAs) for world modeling typically employ fixed-size Vision Transformer encoders that are over-provisioned for simple tasks and under-provisioned for complex ones, with significant redundancy across attention heads. We propose Successive Capacity Growth (SCG), a method that starts from a minimal encoder (1 head, 2 layers, 283K parameters) and grows incrementally in width (adding attention heads for low-level semantic capacity) or depth (adding transformer blocks for higher-order semantic abstraction), driven by a task-agnostic test-and-verify mechanism that exploits function-preserving expansion to safely trial architectural changes and roll back if they do not improve prediction loss. The Sketched Isotropic Gaussian Regularizer (SIGReg) ensures that all learned semantic dimensions remain statistically independent and aligned with the predictive objective, preventing collapse even as the architecture grows. On a 60-dimensional multi-object dynamics task, SCG naturally triggers depth expansion, improving prediction loss by 20.3% over the fixed small baseline with 56 times greater parameter efficiency than scaling to the fixed large model; on a 2D navigation task, a single width expansion yields even an 23% improvement over the fixed large model. Across all three tested environments of increasing complexity, the adaptive encoder matches or exceeds the fixed small baseline, with zero false-positive expansions and bit-exact function preservation (ratio = 1.0, absolute difference = 0.0). The take-away is that JEPA world model encoders need not be pre-allocated at maximum capacity - they can grow successively as the task demands, achieving significant compute and data efficiency while maintaining representation quality.
Comments12 pages, 2 figures, 6 tables