arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Weaver:一种面向AI-RAN计算共享且支持基础模型训练的系统

Weaver: A System for AI-RAN Compute Sharing with Foundation Model Training

Leyang Xue, Tianxin Wang, Xin Zhe Khooi, Jiaxun Yang, Dheeraj Mahendiran, Yufeng Xia, Mun Choon Chan, Myungjin Lee, Mahesh K. Marina

arXiv 2609.35276首次发表:更新:

发表机构

The University of Edinburgh; National University of Singapore; Cisco Research(爱丁堡大学; 新加坡国立大学; 思科研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对AI-RAN基站GPU闲置算力时空分布不均的问题,提出RAN优先的Weaver系统,通过算力感知调度和两级弹性训练框架高效利用闲置算力开展基础模型训练,显著提升可用算力与训练吞吐量。

AI 中文摘要

AI-RAN基础设施为蜂窝基站配备了GPU加速硬件,这一新兴技术为将非RAN工作负载与主RAN处理任务共址部署创造了机遇。我们探索利用这类闲置算力开展基础模型(FMs)的去中心化训练——基础模型训练是计算密集度最高的AI工作负载之一。\n我们首次从两个尺度表征了AI-RAN系统中的GPU闲置算力:微观尺度即单个基站内不同传输时隙的算力情况,宏观尺度即跨基站的算力分布。分析发现,40%-85%的GPU算力处于未使用状态;尽管单个基站的闲置算力在时间上具有突发性,但跨基站的算力在空间上具有互补性。\n为安全高效地利用这些资源,我们提出了Weaver系统,它能在不影响RAN性能的前提下,伺机与对延迟敏感的RAN工作负载并行训练基础模型。Weaver采用RAN优先的设计:集成在MAC调度器中的闲置算力控制器使用算力感知调度来平滑RAN的GPU需求,从而释放更多可用的GPU闲置算力。随后,一个两级弹性训练框架可适配基站内部及跨基站的动态、异构闲置算力。\n在符合O-RAN标准的系统原型上开展的实验表明,Weaver最多可将可用闲置算力提升4.9倍,闲置算力利用率最高可达83%。在多基站测试平台上,与基线方法相比,Weaver将训练吞吐量提升了2.1-3.7倍。

英文摘要

The emergence of AI-RAN infrastructure, which equips cell sites with GPU-accelerated hardware, creates an opportunity to colocate non-RAN workloads with primary RAN processing. We explore using this spare capacity for decentralized training of foundation models (FMs), one of the most compute-intensive AI workloads. We present the first characterization of spare GPU capacity in AI-RAN systems at both micro-scale--across transmission slots within a cell site--and macro-scale--across sites. Our analysis finds that 40-85% of GPU capacity is unused; although this capacity is temporally bursty at individual sites, it is spatially complementary across sites. To safely and efficiently harness these resources, we present Weaver, a system that opportunistically trains FMs alongside latency-critical RAN workloads without degrading RAN performance. Weaver adopts a RAN-first design: a spare-compute controller integrated into the MAC scheduler uses compute-aware scheduling to smooth RAN GPU demand and exposes more usable spare GPU capacity. A two-level elastic training framework then adapts to dynamic, heterogeneous spare capacity within and across sites. Experiments on an O-RAN-aligned system prototype show that Weaver creates up to 4.9x more usable spare compute and utilizes up to 83% of the available spare capacity. On a multi-site testbed, Weaver improves training throughput by 2.1-3.7x over baseline approaches.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑