arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05748cs.DC

MOLT:面向LLM服务的机会式微调细粒度GPU内存共享系统

MOLT: A Fine-Grained GPU Memory Sharing System for LLM Serving with Opportunistic Fine-Tuning

Jaehoon Yang, Yongbeom Kim, Hojoon Kim, Seung Yul Lee, Jae W. Lee

首次发表
浏览论文内容

中文总结 AI 辅助

MOLT提出细粒度内存共享,让推理回收微调步骤中单个激活的内存,反向传播重算,在保持SLO达标率≥99.7%的同时,完成1.9-3.3倍微调工作量。

中文摘要 AI 辅助

大语言模型(LLM)服务根据请求负载扩展其副本数量,但副本内部的GPU内存仍处于空闲状态。添加一个副本需要数分钟,而副本所需的内存在数秒内就会变化。即使即时自动缩放也无法归还这些空闲内存,因为其能移除的最小单元是整个副本。将参数高效微调(PEFT)与推理共同部署可以利用这部分内存,但推理必须能在数秒内回收这些内存,以免等待内存的请求超过其延迟服务级别目标(SLO)。现有的共置系统要么保持微调内存常驻,要么让推理以整个训练样本的粗粒度回收内存。每次这样的回收还会丢弃正在运行的微调步骤。为解决这些限制,我们提出MOLT,一个细粒度内存共享系统,它允许推理回收正在运行的微调步骤为反向传播保存的单个激活的内存。该步骤继续执行,其反向传播会重新计算这些激活。推理仅回收没有在途GPU工作仍可访问的内存,即使在CPU-GPU异步和张量并行下也是如此。在H100 SXM和B200 GPU上的四个模型部署(24B-70B)中,在轨迹驱动的工作负载下,MOLT将推理SLO达标率保持在99.7%或以上,并完成基于丢弃的内存共享1.9-3.3倍的微调工作量。

英文摘要

Large language model (LLM) serving scales its replica count with the request load, yet GPU memory still stands idle inside the replicas. Adding a replica takes minutes, while the memory that a replica needs changes within seconds. Even instant autoscaling could not return this idle memory, because the smallest unit that it can remove is a whole replica. Colocating parameter-efficient fine-tuning (PEFT) with inference can use this memory, but inference must be able to reclaim it within seconds, before requests that wait for memory exceed their latency service-level objective (SLO). Existing colocation systems either keep the tuning memory resident or let inference reclaim it at the coarse granularity of a whole training sample. Each such reclamation also discards the running tuning step. To address these limitations, we present MOLT, a fine-grained memory sharing system that lets inference reclaim the memory of individual activations that a running tuning step has saved for its backward pass. The step continues, and its backward pass recomputes those activations. Inference reclaims only memory that no in-flight GPU work can still access, even under CPU--GPU asynchrony and tensor parallelism. On four model deployments (24B--70B) across H100 SXM and B200 GPUs under trace-driven workloads, MOLT keeps inference SLO attainment at or above 99.7% and completes 1.9--3.3x the tuning work of discard-based memory sharing.

发表机构

  • Seoul National University(首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

↑