发表机构
IIT Guwahati; SustainAI Lab, MFSDS&AI(古瓦哈蒂印度理工学院; SustainAI实验室、MFSDS&AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对稀疏事件驱动视觉Transformer提出参数高效持续学习框架sLoTh,仅更新不足1%参数即可实现无重放的类增量与在线持续学习,能耗约为传统密集Transformer的1/6.5。
AI 中文摘要
机器人与边缘智能系统在动态环境中运行,数据持续流入,要求模型在严格的内存与能耗约束下适应新任务的同时保留已学知识。尽管参数高效微调在视觉Transformer的持续学习中展现出潜力,但传统架构依赖密集计算,实际部署成本高昂。稀疏事件驱动视觉Transformer提供高能效的事件驱动计算,但其持续学习能力仍未被充分探索。本文提出sLoTh,一种针对预训练稀疏事件驱动(脉冲)视觉Transformer的参数高效持续学习框架。sLoTh冻结主干网络,将可塑性限制于可扩展高效低秩注意力更新(seLoRA)与共享神经元阈值调制,无需重放缓冲区即可实现适应,仅更新模型不到1%的参数。在CIFAR-100、Tiny-ImageNet、ImageNet-100及ImageNet-R数据集上,针对最多100个任务的实验表明,该方法在类增量学习与在线持续学习中具备具竞争力的无重放性能,同时能耗约为传统密集视觉Transformer的1/6.5。
英文摘要
Robotic and edge intelligence systems operate in dynamic environments where data arrives continuously, requiring models to adapt while preserving previously learned knowledge under strict memory and energy constraints. While parameter-efficient fine-tuning has shown promise for continual learning with vision transformers, conventional architectures rely on dense computation and remain costly for real-world deployment. Sparse event-based vision transformers provide energy-efficient event-driven computation, yet their continual learning capabilities remain largely unexplored. We here introduce sLoTh, a parameter-efficient continual learning framework for pretrained sparse event-based (spiking) vision transformers. sLoTh freezes the backbone and restricts plasticity to scalable-efficient low-rank attention updates (seLoRA) and shared neuronal threshold modulation, enabling adaptation without replay buffers by updating less than 1% of model parameters. Experiments across CIFAR-100, Tiny-ImageNet, ImageNet-100, and ImageNet-R with up to 100 tasks demonstrate competitive rehearsal-free performance in class-incremental learning and online continual learning, while enabling approximately 6.5x lower energy consumption than conventional dense vision transformers.