面向机器学习的机器学习
ML-for-ML
- Politecnico di Milano(米兰理工大学)
- Broadcom(博科公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对AI训练工作负载的网络与ML系统分离优化导致性能浪费问题,提出跨层联合优化的ML-for-ML方案,可使达到目标损失的速度提升最多42%。
AI中文摘要:
AI训练工作负载正快速增长,其时间、能耗和基础设施成本愈发重要。在共享云集群中,训练与微调作业会与协同运行的工作负载竞争网络资源,而网络机制与ML训练选择通常是分别优化的:网络控制字节的传输方式,ML系统控制通信的时机与规模。我们认为这种分离会导致端到端性能被浪费。我们提出ML-for-ML,这是一种跨层视角,在共享的达到目标损失时间目标下,联合选择网络侧与ML侧的旋钮参数。我们的初步原型显示,通过协同优化ML和网络参数,达到目标损失的速度可提升多达42%。
英文摘要:
AI training workloads are growing rapidly, making their time, energy, and infrastructure costs increasingly important. In shared cloud clusters, training and fine-tuning jobs compete with co-running workloads for network resources, while network mechanisms and ML training choices are typically optimized separately: networking controls how bytes move, whereas ML systems control when and how much communication occurs. We argue that this separation leaves end-to-end performance on the table. We present ML-for-ML, a cross-layer perspective in which network-side and ML-side knobs are selected jointly under a shared time-to-target-loss objective. Our preliminary prototype shows that by co-optimizing the ML and network parameters, we reach the target loss up to 42% faster.