arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33325cs.CV

VisionHOPE:视觉骨干网络作为自修改学习系统

VisionHOPE: Visual Backbones as Self-Modifying Learning Systems

Siran Peng, Tianshuo Zhang, Tianyu Fu, Weisong Zhao, Haoyuan Zhang, Jiankuo Zhao, Minghui Wu, Ping Jiang, Xiangyu Zhu, Chenxu Zhao, Zhen Lei

首次发表
浏览论文内容

中文总结 AI 辅助

VisionHOPE提出首个自修改学习系统的视觉骨干网络,通过五个耦合记忆实现内容与学习方式的共同演化,并采用稳定性匹配的步长控制,在ImageNet-1K、COCO和ADE20K上取得竞争性结果。

中文摘要 AI 辅助

视觉骨干网络已经从具有局部聚合的卷积神经网络(CNN)演进到具有全局交互的视觉Transformer(ViT),再到具有输入相关状态转移的状态空间模型(SSM),以及在处理图像时适应内部学习器的测试时训练(TTT)层。在这一演进过程中,视觉计算对每个输入的适应性越来越强,但控制这种适应的规则在很大程度上仍由训练好的骨干网络预设。我们提出了VisionHOPE,这是第一个将视觉骨干网络构建为自修改学习系统的通用框架,在该系统中,模型记忆的内容以及学习的方式在图像内部共同演化。基于嵌套学习(NL)的自引用构造,VisionHOPE通过五个耦合的记忆实现这种共同演化,这些记忆存储内容、生成键和值表示,并控制学习率和保留率。随着视觉上下文沿每次扫描累积,这些记忆共同演化。然而,直接将无约束的自引用更新应用于视觉骨干网络会导致不稳定。因此,我们推导了一种稳定性匹配的步长控制方案,该方案将自引用注入的软上限与保留记忆转换的光谱钳制相结合,并证明了所得记忆动力学沿每次扫描是非扩张的。对于二维特征图,我们通过将块与图像行和列对齐,在四个方向扫描中调整NL的块公式。所提出的VisionHOPE在ImageNet-1K、COCO和ADE20K上取得了有竞争力的结果,将自修改学习系统确立为通用视觉骨干网络的实用基础。代码可在该https URL获取。

英文摘要

Visual backbones have evolved from Convolutional Neural Networks (CNNs) with local aggregation to Vision Transformers (ViTs) with global interactions, State-Space Models (SSMs) with input-dependent state transitions, and Test-Time Training (TTT) layers that adapt an inner learner while processing an image. Across this progression, visual computation has become increasingly adaptive to each input, yet the rules governing that adaptation remain largely prescribed by the trained backbone. We introduce VisionHOPE, the first generic visual backbone formulated as a self-modifying learning system, in which what the model remembers and how it learns co-evolve within an image. Building on the self-referential construction of Nested Learning (NL), VisionHOPE realizes this co-evolution through five coupled memories that store content, generate key and value representations, and govern learning rate and retention. These memories evolve jointly as visual context accumulates along each scan. However, directly applying the unconstrained self-referential update to a visual backbone leads to instability. We therefore derive a stability-matched step-size control scheme that combines a soft cap on self-referential injection with a spectral clamp on the retained memory transition, and prove that the resulting memory dynamics are non-expansive along each scan. For two-dimensional feature maps, we adapt NL's chunk formulation by aligning chunks with image rows and columns across four directional scans. The proposed VisionHOPE achieves competitive results on ImageNet-1K, COCO, and ADE20K, establishing self-modifying learning systems as a practical foundation for general-purpose visual backbones. The code is available at https://github.com/PSRben/VisionHOPE.

↑