arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06046cs.NIcs.DCcs.LG

面向机器学习的机器学习

ML-for-ML

  • Politecnico di Milano(米兰理工大学)
  • Broadcom(博科公司)

机构由 AI 辅助整理,请以论文原文为准。

Yutong Zhao, Noga H. Rotman, Gianni Antichi, Ran Ben Basat

AI总结:

针对AI训练工作负载的网络与ML系统分离优化导致性能浪费问题,提出跨层联合优化的ML-for-ML方案,可使达到目标损失的速度提升最多42%。

AI中文摘要:

AI训练工作负载正快速增长,其时间、能耗和基础设施成本愈发重要。在共享云集群中,训练与微调作业会与协同运行的工作负载竞争网络资源,而网络机制与ML训练选择通常是分别优化的:网络控制字节的传输方式,ML系统控制通信的时机与规模。我们认为这种分离会导致端到端性能被浪费。我们提出ML-for-ML,这是一种跨层视角,在共享的达到目标损失时间目标下,联合选择网络侧与ML侧的旋钮参数。我们的初步原型显示,通过协同优化ML和网络参数,达到目标损失的速度可提升多达42%。

英文摘要:

AI training workloads are growing rapidly, making their time, energy, and infrastructure costs increasingly important. In shared cloud clusters, training and fine-tuning jobs compete with co-running workloads for network resources, while network mechanisms and ML training choices are typically optimized separately: networking controls how bytes move, whereas ML systems control when and how much communication occurs. We argue that this separation leaves end-to-end performance on the table. We present ML-for-ML, a cross-layer perspective in which network-side and ML-side knobs are selected jointly under a shared time-to-target-loss objective. Our preliminary prototype shows that by co-optimizing the ML and network parameters, we reach the target loss up to 42% faster.

补充信息

↑