发表机构
Meta Platforms, Inc.(Meta平台公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文介绍Meta自研多主机NIC fbnic的运营基础设施,通过物理隔离、硬件在环CI、统一观测和限定修复,实现计划外不可用性降低12倍、修复时间降低37%。
AI 中文摘要
我们描述了为在Meta的数十万台生产主机上部署和运营fbnic(一种自定义多主机NIC)而构建的运营基础设施。供应商的多主机NIC,通过改造单主机架构设计,存在共享固件和缓冲区的问题,在七年中导致了级联的隔离故障。fbnic通过物理隔离消除了这些问题,但转向自研硬件将整个运营负担转移到了超大规模数据中心运营商身上。我们提出了一个硬件在环的持续集成流水线,用于测试固件、驱动程序和内核的交叉产品;一个统一的观测流水线,将NIC和交换机计数器共同定位,用于跨层故障归因;一个驱动优先的架构,固件消息类型少于十种;一个针对性的固件升级编排器,粒度低于机架;以及限定的修复自动化,将爆炸半径限制在单个主机切片内。在十个月的时间里,与同一平台上的供应商NIC相比,fbnic实现了计划外不可用性降低12倍,平均修复时间降低37%,硬件更换次数减少2.3倍。
英文摘要
We describe the operational infrastructure built to deploy and operate fbnic, a custom multi-host NIC, across hundreds of thousands of production hosts at Meta. Vendor multi-host NICs, designed by retrofitting single-host architectures, suffered from shared firmware and buffers that created cascading isolation failures over seven years. fbnic eliminates these through physical isolation, but shifting to in-house hardware shifts the entire operational burden to the hyperscaler. We present a hardware-in-the-loop CI pipeline testing firmware, driver, and kernel cross-products; a unified observability pipeline co-locating NIC and switch counters for cross-layer fault attribution; a driver-first architecture with fewer than ten firmware message types; a targeted firmware upgrade orchestrator at sub-sled granularity; and scoped repair automation confining blast radius to individual host slices. Over ten months, fbnic achieved a 12X reduction in unplanned unavailability, 37% lower mean time to repair, and 2.3X fewer hardware swaps compared to vendor NICs on the same platform.
Comments15 pages, 7 figures, to be published in NSDI'27 - 24th USENIX Symposium on Networked Systems Design and Implementation