可编程数据平面交换机中优先级调度的并行架构
Parallel Architectures For Priority Scheduling In Programmable Data Plane Switches
- Indian Institute of Technology Madras(印度马德拉斯理工学院)
- Ciena Corporation(思科恩公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对可编程数据平面交换机中优先级调度可扩展性差和近似调度器正确性低的问题,提出P3PO和PPQA两种并行优先级队列架构,将全局队列分解为多个并行子队列,在保持排序的同时提升可扩展性,并通过仿真验证了其有效性。
AI中文摘要:
可编程数据平面(PDP)交换机支持灵活的数据包处理,但在调度能力方面仍然受限,尤其是对于需要根据用户定义的等级严格排序数据包的基于优先级的策略。早期工作称为先进先出(Push-in First-out, PIFO),提出了一种优先级队列,能够根据优先级对入队的数据包进行排序。它为可编程数据包调度提供了一种理想的抽象;然而,其对线速数据包排序的要求使其在高网络速度(100 Gbps及以上)下不切实际。诸如SPPIFO、AIFO和RIFO等近似调度器降低了复杂度,但引入了优先级反转,从而恶化了延迟、公平性和流完成时间(FCTs),因为它们依赖于先进先出(FIFO)队列或使用主动队列管理(AQM)作为调度的替代。为了解决精确调度器的可扩展性限制,同时避免近似调度器的正确性问题,本工作提出了两种数据包排队架构:并行处理优先级排序(P3PO)和并行PIFO排队架构(PPQA)。两种架构都将一个大的全局优先级队列分解为多个并行运行的小型有界容量优先级队列。P3PO使用一组级联的优先级队列,并带有基于优先级的界限,以确保数据包被插入到能够维持排序的最早队列中。PPQA根据占用率将数据包分布到多个独立的优先级队列中,使用解复用器实现均衡排队,使用复用器在出队时全局选择最高优先级的数据包。这些架构在NetBench数据包级模拟器中实现,并使用经验数据中心工作负载(Web搜索和数据挖掘)进行评估。
英文摘要:
Programmable Data Plane (PDP) switches enable flexible packet processing but remain limited in scheduling capabilities, particularly for priority-based policies that require strict ordering of packets according to user-defined ranks. An earlier work, called Push-in First-out (PIFO), proposed a priority queue that enables ordering enqueued packets based on their priority. It provided an ideal abstraction for programmable packet scheduling; however, its requirement for line-rate packet sorting makes it impractical at high network speeds (100 Gbps and beyond). Approximate schedulers such as SPPIFO, AIFO, and RIFO reduce complexity but introduce priority inversions, thus degrading latency, fairness, and Flow completion times (FCTs) since they rely on First-in First-out (FIFO) queues or use Active Queue Management (AQM) as a substitute for scheduling. To address the scalability limitations of accurate schedulers while avoiding the correctness issues of approximate schedulers, this work proposes two packet queuing architectures, Parallel Processing for Priority Ordering (P3PO) and Parallel PIFO Queuing Architecture (PPQA). Both architectures decompose a large global priority queue into multiple smaller bounded-capacity priority queues that operate in parallel. P3PO uses a cascading set of priority queues with priority-based bounds to ensure that packets are inserted into the earliest queue capable of maintaining ordering. PPQA distributes packets across multiple independent priority queues based on occupancy, using a demultiplexer for balanced queuing and a multiplexer for globally selecting the highest-priority packet at dequeue. The architectures are implemented in the NetBench packet-level simulator and evaluated using empirical datacenter workloads (Web-search and Data-mining).