arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MemFLoRA:面向边缘CNN适配的内存下限LoRA

MemFLoRA: Memory-Floor LoRA for CNN Adaptation at the Edge

Mehmet Emre Akbulut, Johannes Geier, Ulf Schlichtmann

arXiv 2610.08669首次发表:更新:

发表机构

Technical University of Munich(慕尼黑工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MemFLoRA提出内存优先的低秩CNN适配器,通过冻结下投影、训练上投影并采用激活最小化反向规则,在HAR任务上减少98.5%以上激活内存,同时保持或超越PEFT基线性能。

AI 中文摘要

当模型在部署后遇到用户、传感器或环境特定的偏移时,设备端学习是必要的。尽管参数高效微调(PEFT)方法,特别是低秩适配(LoRA)变体,能够在边缘实现高效适配,但卷积神经网络(CNN)适配的瓶颈资源往往不是可训练参数的数量,而是必须保留至反向传播的激活状态。本文介绍了内存下限LoRA(MemFLoRA),这是一种基于内存优先设计原则而非直接应用面向Transformer的LoRA的低秩CNN适配器。我们不仅减少可训练权重,还定义了一个激活内存下限标准:可训练的反向计算不得依赖于全宽度层输入。由此产生的适配器冻结下投影,训练尺度匹配的上投影,并将评估模式的主干归一化与激活最小化的反向规则相结合,将保存的状态减少到低秩分支。在三个人类活动识别(HAR)数据集和两个CNN主干上,针对主体、身体位置和传感器放置偏移进行评估,与全微调相比,MemFLoRA将保存的激活内存减少了98.5-98.7%,峰值训练状态内存减少了94.9-97.3%,同时匹配或超过了CNN PEFT基线。

英文摘要

On-device learning is necessary when the model encounters user-,sensor-, or environment-specific shifts after deployment. Although parameter-efficient fine-tuning (PEFT) methods, particularly Low-Rank Adaptation (LoRA) variants, enable efficient adaptation at the edge, the limiting resource for Convolutional Neural Network (CNN) adaptation is often not the number of trainable parameters but the activation state that must be retained until the backward pass. This paper introduces Memory-Floor LoRA (MemFLoRA), a low-rank CNN adapter built around a memory-first design principle rather than a direct application of transformer-oriented LoRA. Instead of merely reducing trainable weights, we define an activation-memory-floor criterion: trainable backward computations must not depend on full-width layer inputs. The resulting adapter freezes the down-projection, trains a scale-matched up-projection, and combines eval-mode backbone normalization with activation-minimal backward rules, reducing saved state to the low-rank branch. Evaluated on three Human Activity Recognition (HAR) datasets and two CNN backbones under subject, body-location, and sensor-placement shifts, MemFLoRA reduces saved-activation memory by 98.5-98.7% and peak training-state memory by 94.9-97.3% relative to full fine-tuning, while matching or exceeding CNN PEFT baselines.

CommentsAccepted at the 32nd Asia and South Pacific Design Automation Conference (ASP-DAC 2027), January 25-28, 2027, Tokyo, Japan. Code: https://github.com/mehmetemreakbulut/MemFLoRA

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑