arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AutoUVM:UVM 超订下面向大语言模型的自动化预取框架

AutoUVM: Automated Prefetching Framework for LLMs under UVM Oversubscription

Mao Lin, Hui Feng, Xianzhong Ding, Guilherme Cox, Qian Wang, Hyeran Jeon

arXiv 2609.06172首次发表:更新:

发表机构

University of California, Merced; NVIDIA(加州大学默塞德分校; 英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AutoUVM 提出自动化框架感知的 UVM 预取系统,通过张量级细粒度预取策略,在内存超订下平均加速 LLM 执行 3.1 倍,超越现有预取器。

AI 中文摘要

大语言模型(LLM)日益超出商用 GPU 的内存容量,使得内存超订在实际部署中变得普遍。NVIDIA 统一虚拟内存(UVM)提供了对主机内存的透明访问,但其由缺页驱动的迁移引入了严重的性能开销。尽管 UVM 暴露了原语(例如预取和放置提示)来缓解这些成本,但它们需要低层 CUDA 修改,限制了其对大多数 LLM 用户的适用性。同时,现有的 UVM 优化在粗粒度的托管对象粒度上操作,未能捕获深度学习框架内部的张量级内存行为,导致过多的数据移动和 CPU-GPU 互连瓶颈。我们提出了 AutoUVM,一个自动化的、框架感知的 UVM 预取系统,用于在内存超订下高效执行 LLM。AutoUVM 通过暴露张量级访问信息并以细粒度实现策略驱动的预取,弥合了深度学习框架与 UVM 之间的语义鸿沟。作为透明扩展实现,AutoUVM 无需更改模型代码,并动态适应运行时内存压力。我们使用受屋顶线(roofline)启发的策略实例化 AutoUVM,以识别性能关键的数据传输。在十个 LLM 上,AutoUVM 相比基线 UVM 实现了平均 3.1 倍的加速,并持续超越先前最佳 UVM 预取器 1.9 倍,相比对象级预取器最高提升 4.7 倍,同时显著减少了缺页。

英文摘要

Large language models (LLMs) increasingly exceed the memory capacity of commodity GPUs, making memory oversubscription common in practical deployments. NVIDIA Unified Virtual Memory (UVM) provides transparent access to host memory, but its page-fault-driven migrations introduce severe performance overhead. While UVM exposes primitives (e.g., prefetching and placement hints) to mitigate these costs, they require low-level CUDA modifications, limiting their applicability for most LLM users. Meanwhile, existing UVM optimizations operate at coarse managed-object granularity and fail to capture deep learning frameworks' internal tensor-level memory behavior, leading to excessive data movement and CPU-GPU interconnect bottlenecks. We propose AutoUVM, an automated, framework-aware UVM prefetching system for efficient LLM execution under memory oversubscription. AutoUVM bridges the semantic gap between deep learning frameworks and UVM by exposing tensor-level access information and enabling policy-driven prefetching at fine granularity. Implemented as a transparent extension, AutoUVM requires no changes to model code and dynamically adapts to runtime memory pressure. We instantiate AutoUVM with a roofline-inspired policy to identify performance-critical data transfers. Across ten LLMs, AutoUVM achieves an average 3.1x speedup over baseline UVM and consistently surpasses the best-performing prior UVM prefetcher by 1.9x, with improvements of up to 4.7x over object-level prefetchers, while significantly reducing page faults.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑