arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21288cs.AR

评估三模冗余用于宽链路低延迟NoC路由器:可靠性与物理设计挑战

Assessing Triple Modular Redundancy for Wide-Link, Low-Latency NoC Routers: Reliability and Physical Design Challenges

Chen Wu, Michael Rogenmoser, Luca Benini, Angelo Garofalo

首次发表
浏览论文内容

中文总结 AI 辅助

本文评估三种粒度的三模冗余(TMR)用于512位宽链路2周期延迟NoC路由器,发现全TMR可消除超百万次注入故障,集成至AI加速tile后仅增16.8%面积,为物理AI系统NoC提供实用防护方案。

中文摘要 AI 辅助

在恶劣环境中部署的基于物理AI tile的加速器,其片上网络(Network-on-Chip, NoC)需防护单粒子效应(Single-Event Effects, SEEs),这对避免NoC故障(可能导致死锁和静默数据损坏(Silent Data Corruption, SDC))至关重要。以往可靠NoC的研究多聚焦于窄链路(如32位)、深度流水线路由器及单粒子翻转(Single-Event Upsets, SEUs),但当前技术已演进为采用先进工艺节点、工作频率超1 GHz的低延迟NoC路由器,且具备超宽链路。本文针对一款带512位宽链路、延迟为2周期的NoC路由器,评估三种粒度(粗粒度、仅状态、全粒度)的三模冗余(Triple Modular Redundancy, TMR)实现的成本与可靠性权衡。我们采用台积电(TSMC)7nm工艺完成从寄存器传输级(RTL)到图形数据库系统接口(GDSII)的物理设计,并开展RTL级和网表级的SEU与单粒子瞬态(Single-Event Transients, SET)故障注入实验。我们从可靠性、成本及物理设计策略三方面评估这三种TMR方案,进一步将评估范围从独立路由器扩展至完整AI加速tile。结果表明,仅状态TMR与粗粒度TMR无法提供足够的SEEs防护,而全TMR在每次实验中注入超100万个故障的情况下,可消除所有观测到的故障。尽管独立的全TMR路由器会带来7.04倍的面积开销,但将其集成至包含处理器和本地一级缓存(L1 memories)的完整AI加速tile后,该成本会大幅分摊:在系统级的通用矩阵乘(GEMM)基准测试中,相同设计仅增加16.8%的面积和15.2%的功耗,且tile的关键路径完全未受影响。这些结果表明,先进工艺节点提供了充足的布线容量,使全TMR成为在恶劣环境中运行的物理AI系统中防护NoC的实用且可部署的解决方案。

英文摘要

Protecting the Network-on-Chip (NoC) of physical-AI tile-based accelerators deployed in harsh environments against single-event effects (SEEs) is paramount for preventing NoC failures that can lead to deadlocks and silent data corruption (SDC). Prior work on reliable NoCs has largely focused on narrow links (e.g., 32-bit), deeply pipelined routers, and single-event upsets (SEUs). However, the state of the art has evolved toward low-latency NoC routers with ultra-wide links, implemented on advanced technology nodes and operating at frequencies above 1 GHz. We evaluate the cost and reliability trade-offs of implementing Triple Modular Redundancy (TMR) at three granularities (coarse, state-only, and full) for a 2-cycle-latency NoC router with 512-bit wide links. We carry out RTL-to-GDSII physical design in TSMC 7nm technology, as well as both RTL- and netlist-level SEU and SET fault injection campaigns. We evaluate the three TMR approaches in terms of reliability, cost, and physical design strategies, further extending the assessment from a standalone router to a full AI acceleration tile. Our results show that state-only and coarse-grained TMR do not provide sufficient protection against SEEs, whereas full TMR eliminates all observed failures across more than one million injected faults per experiment. Although the standalone full-TMR router incurs a 7.04x area overhead, this cost is drastically amortized once integrated into a complete AI accelerator tile with processors and local L1 memories: the same design adds only 16.8% area and 15.2% power consumption under a GEMM benchmark at the system level, with the critical path of the tile entirely unaffected. These results demonstrate that advanced technology nodes provide sufficient routing capacity to make full TMR a practical and deployable solution for protecting NoCs in Physical AI systems operating in harsh environments.

补充信息

↑