AI 中文总结
ProFlow是一种基于强化学习的主动式数据中心流放置框架,可利用网络拥塞早期征兆提前重路由受保护流,相比反应式方法吞吐量提升约40%且决策提前约34秒。
AI 中文摘要
在由叶交换机和汇聚交换机组成的数据中心架构中,竞争流可能会在共享汇聚交换机上共置,引发拥塞,大幅降低受保护流的性能。然而,在吞吐量下降变得可观测之前,网络通常会出现早期征兆,表现为流活动上升和队列溢出信号。现有拥塞管理方法主要在拥塞显现后才做出反应,未充分利用这些早期征兆。本文提出ProFlow,这是一种用于多租户数据中心网络中保护性能敏感流量的主动式流放置框架,从而利用潜在吞吐量下降的早期征兆。ProFlow利用分布式遥测信号和离线训练的强化学习(RL)来识别前驱拥塞状况,并在吞吐量下降发生前主动重路由受保护流。使用FABRIC测试平台的评估结果显示,ProFlow的平均吞吐量比反应式重路由基线高约40%,且平均在约34秒前启动重路由决策,证明了 anticipatory 拥塞管理的有效性。
英文摘要
In datacenter fabrics composed of leaf and aggregation switches, competing flows may become co-located on shared aggregation switches, creating congestion that can significantly degrade protected flows. However, before throughput degradation becomes observable, the network often exhibits early signs characterized by rising flow activity and queue overflow signals. Existing congestion-management approaches primarily react only after congestion becomes visible, leaving these early signs largely unexploited. In this paper, we propose ProFlow, a proactive flow-placement framework for protecting performance-sensitive traffic in multi-tenant datacenter networks, thereby utilizing the early signs of potential throughput degradations. ProFlow leverages distributed telemetry signals and offline-trained reinforcement learning (RL) to identify precursor congestion conditions and proactively reroute protected flows before throughput degradation occurs. Evaluation results using FABRIC testbed show that ProFlow achieves approximately 40% higher mean throughput than a reactive rerouting baseline while initiating rerouting decisions around 34 seconds earlier on average, demonstrating the effectiveness of anticipatory congestion management.
Comments13 Pages