arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

间接寻址的力量:将交换机扩展到硅边界之外

The Power of Indirection: Scaling Switches Beyond Silicon Boundaries

Lukas Röllin, Sushovan Das, Paolo Costa, Laurent Vanbever

arXiv 2609.31092首次发表:更新:

发表机构

ETH Zürich; Microsoft Research(苏黎世联邦理工学院; 微软研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多ASIC交换机中ASIC间带宽瓶颈问题,提出引入电路交换间接层的Fastroute架构,结合分组与电路交换,在降低带宽需求的同时实现接近单片ASIC的性能,并通过硬件原型验证其可行性。

AI 中文摘要

摩尔定律的放缓以及单片集成的面积限制,使得基于小芯片(chiplet)的设计在许多领域(包括网络ASIC)中成为必然。然而,将多个网络ASIC组合在一起面临一个根本性挑战:维持足够的ASIC间带宽以匹配理想化单片ASIC设计的性能。提供全带宽代价高昂,因为它需要宝贵的转发容量,而减少ASIC间带宽则会造成严重的性能瓶颈。我们提出了一种新颖的多ASIC交换机架构,在ASIC前端引入一个电路交换的间接层(indirection layer)。该层灵活地将入口端口重新映射到各个ASIC,根据观察到的流量模式实现流量本地化并最小化ASIC间通信。我们的系统Fastroute结合了分组交换和电路交换,在降低ASIC间带宽需求的同时,提供与单片ASIC交换机相当的性能。这释放了用于外部网络接口的容量,使Fastroute能够超越传统的非过订阅(non-oversubscribed)多ASIC设计。我们的硬件原型通过在大语言模型(LLM)训练工作负载上的评估,证明了该系统的功能可行性。通过降低带宽和功耗开销,Fastroute弥合了硅制造限制与日益增长的应用程序需求之间的差距。它为向多ASIC交换机的过渡提供了一条高效路径,使得在不等待下一代ASIC的情况下即可满足带宽和端口数(radix)需求。

英文摘要

The slowdown of Moore's law and the area limit of monolithic integration have made chiplet-based designs inevitable across many domains, including network ASICs. However, combining multiple network ASICs together poses a fundamental challenge: maintaining sufficient inter-ASIC bandwidth to match the performance of an idealistic single-ASIC design. Providing full bandwidth is prohibitively expensive as it requires valuable forwarding capacity, while reducing inter-ASIC bandwidth creates severe performance bottlenecks. We propose a novel multi-ASIC switch architecture that introduces a circuit-switched indirection layer in front of the ASICs. This layer flexibly remaps ingress ports across ASICs, localizing traffic and minimizing inter-ASIC communication based on observed patterns. Our system, Fastroute, combines packet and circuit switching to deliver performance comparable to a single-ASIC switch while reducing inter-ASIC bandwidth requirements. This frees up capacity for external network interfaces, allowing Fastroute to outperform traditional non-oversubscribed multi-ASIC designs. Our hardware prototype demonstrates the system's functional feasibility by evaluating it on an LLM training workload. By reducing bandwidth and power overhead, Fastroute bridges the gap between silicon fabrication limits and soaring application demands. It provides an efficient transition to multi-ASIC switches, enabling bandwidth and radix demand to be met without waiting for the next ASIC generation.

Comments19 pages, 26 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑