发表机构
University of Macau; University of Georgia; Hong Kong University of Science and Technology(澳门大学; 佐治亚大学; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对移动GPU上Transformer模型微调内存低效问题,提出FBLayout框架,通过统一布局、索引转换及激活引导选择等方法,在多手机多模型测试中实现加速,提高缓存效率并减少内存占用,助力设备上大模型微调。
AI 中文摘要
基于Transformer的模型在语言、视觉和多模态任务中展现出前所未有的能力。在设备上对Transformer模型进行微调为个性化AI提供了隐私保护途径,但由于训练期间严重的内存限制和注意力机制中频繁的布局转换,在移动GPU上效率仍然低下。现有的移动训练框架要么在前向和反向传播中使用统一布局,导致反向传播期间内存访问碎片化和GPU利用率低,要么依赖显式布局转换,引入大量转换开销。为克服此问题,我们提出FBLayout,一种与移动GPU平台协同设计张量组织的布局感知框架。FBLayout引入:(1)用于前向/反向传播中多维归约的统一R-Tile布局;(2)基于tile的索引转换以消除物理数据移动;(3)激活引导的布局选择以全局传播高效布局。在不同手机(包括ARM Mali和高通Adreno GPU)上对七个Transformer模型的评估表明,FBLayout比MNN、TFLite和TVM实现了2.2-5.7倍的加速,同时显著提高缓存效率并减少内存占用,实现了实用的设备上大模型微调。
英文摘要
Transformer-based models have enabled unprecedented capabilities across language, vision, and multimodal tasks. On-device fine-tuning of transformer models offers a privacy-preserving path to personalized AI, yet remains inefficient on mobile GPUs due to severe memory constraints and frequent layout transformations in attention mechanism during training. Existing mobile training frameworks either use unified layouts for forward and backward passes -- leading to fragmented memory access and poor GPU utilization during backpropagation -- or rely on explicit layout conversions, which introduce significant transformation overhead. To overcome this, we propose FBLayout, a layout-aware framework that co-designs tensor organization with mobile GPU platforms. FBLayout introduces: (1) a unified R-Tile layout for multi-dimensional reductions across forward/backward passes; (2) tile-based index transformation to eliminate physical data movement; and (3) activation-guided layout selection to propagate efficient layouts globally. Evaluations on seven transformer models across different mobile phones (including ARM Mali and Qualcomm Adreno GPUs) show that FBLayout achieves 2.2-5.7x speedup over MNN, TFLite, and TVM, while significantly improving cache efficiency and reducing memory footprint, enabling practical on-device large model fine-tuning.
Comments13 pages, 21 figures, Mobisys 2026