一次驻留,轻量交换:面向高效Ultra-LoRA服务的子空间对齐质心-残差训练
Pin Once, Swap Light: Subspace-Aligned Centroid-Residual Training for Efficient Ultra-LoRA Serving
浏览论文内容
中文总结 AI 辅助
针对多租户LoRA服务中高秩适配器开销大、超低秩适配器性能差的困境,提出分层微调框架SALT,通过对齐领域质心与超低秩残差,兼顾性能与服务效率。
中文摘要 AI 辅助
现代多租户低秩适配器(Low-Rank Adapters, LoRA)服务系统会同时托管数十到数百个LoRA适配器。这类系统虽功能强大,却在服务效率与任务性能之间陷入关键的系统困境:更高秩的适配器通常能取得更优的下游任务性能,但其GPU显存占用以及主机到设备的PCIe交换开销会严重限制可扩展性。反之,超低秩适配器($r \le 2$)能将显存占用与PCIe传输开销降至最低,却会出现下游任务性能下降的问题。为解决该问题,我们提出子空间对齐LoRA训练(Subspace-Aligned LoRA Training, SALT)——一种具备服务效率感知的分层微调框架。我们的方案分为三个阶段:首先,服务提供方在领域内的公开数据上联合训练高容量领域质心,采用一种新型对齐正则化器将域内任务子空间凝聚为统一基;其次,用户在这些冻结的质心之上,基于私有数据微调超低秩任务残差适配器;最后,推理阶段服务提供方将质心驻留在GPU显存中,并根据需求动态换入每个用户的任务残差。在不同规模的大语言模型(LLM)上,SALT使用$r \le 2$的残差即可恢复高秩精度,相比现有最优压缩基线,绝对精度提升最高达18.5%,单适配器内存占用最高降低16倍。将SALT集成到vLLM后,针对Llama-3.2-3B模型,在PCIe带宽受限场景下服务吞吐量最高提升51%,在GPU显存受限场景下最高提升28%。
英文摘要
Modern multi-tenant Low-Rank Adapters (LoRAs) serving systems concurrently host tens to hundreds of LoRA adapters. Though powerful, this introduces a critical system dilemma between serving efficiency and task performance: higher-rank adapters generally achieve better downstream task performance, but their GPU VRAM footprint and Host-to-Device PCIe swapping overhead severely constrain scalability. Conversely, ultra-low-rank adapters ($r \le 2$) minimize both VRAM footprint and PCIe transfer overhead, but suffer from downstream task performance degradation. To solve this problem, we propose Subspace-Aligned LoRA Training (SALT), a serving efficiency-aware hierarchical fine-tuning framework. Our solution operates in three phases. First, a provider jointly trains high-capacity domain centroids on public data within the domain using a novel alignment regularizer that coheres in-domain task subspaces into a unified basis. Next, users fine-tune ultra-low-rank task residual adapters on private data atop those frozen centroids. Finally, during inference, the provider pins the centroid in GPU VRAM and dynamically swaps in each user's task residual on demand. Across LLMs of varying scales, SALT recovers high-rank accuracy using $r \le 2$ residuals, achieving up to 18.5% absolute accuracy gains over state-of-the-art compression baselines and reducing per-adapter memory by up to 16x. When integrated into vLLM, SALT improves serving throughput by up to 51% under PCIe bandwidth pressure and 28% under GPU VRAM constraints for Llama-3.2-3B.