发表机构
University of Bologna; ETH Zurich; lowRISC C.I.C.; Tenstorrent USA, Inc.(博洛尼亚大学; 苏黎世联邦理工学院; lowRISC 公益公司; Tenstorrent 美国公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对混合关键度SoC的AXI4协议缺陷,提出PLT、CLT、ILT三种监控方案,经实验验证可100%检测相关故障,CLT与ILT分别实现可观面积缩减。
AI 中文摘要
基于AXI4开放标准协议的带片上互连的混合关键度片上系统(SoC)缺乏协议级超时机制,当下属设备或管理器因硬件故障、辐射诱发的翻转或软件错误而失效或停滞时,系统会面临死锁和错过实时时限的风险。本研究提出了一种可配置的硬件知识产权(IP),在无故障运行时非侵入式工作,可在运行时检测AXI4协议违规和时序故障,并通过“切断-排空”隔离机制恢复互连活性。为解决监控粒度与面积开销之间的根本权衡,我们提出了三种监控粒度递减的设计:阶段级跟踪(PLT),可在单个协议阶段提供周期精确的故障定位;通道级跟踪(CLT),将各阶段监控器合并为通道级监督;ID级跟踪(ILT),仅通过监控每个ID的事务边界实现亚线性面积扩展。采用GlobalFoundries 12nm工艺综合后,CLT相较PLT减少了36.7%的面积,同时保持了最坏情况检测界限和极低的检测延迟开销;而ILT实现了89.2%的面积缩减,适用于严格受限的部署,代价是检测延迟中位数升高3.7倍,故障定位粒度更粗。对RISC-V SoC开展的120万场景故障注入实验证实,所有表现为AXI4协议或活性违规的故障均未被漏检,观测到的检测延迟始终受限于理论最坏情况预测。
英文摘要
Mixed-criticality Systems-on-Chip (SoCs) with on-chip interconnects based on the AXI4 open standard protocol lack a protocol-level timeout mechanism, exposing systems to deadlocks and missed real-time deadlines when subordinate devices or managers fail or stall due to hardware faults, radiation-induced upsets, or software errors. This work presents a configurable hardware intellectual property (IP), non-intrusive in fault-free operation, that detects AXI4 protocol violations and timing faults at runtime and restores interconnect liveness through a cut-and- drain isolation mechanism. To address the fundamental trade-off between monitoring granularity and area cost, we introduce three designs at decreasing monitoring granularity: Phase-Level Track-ing (PLT), which provides cycle-accurate fault localization across individual protocol phases; Channel-Level Tracking (CLT), which coalesces per-phase monitors into channel-level supervision; and ID-Level Tracking (ILT), which achieves sub-linear area scaling by monitoring only per-ID transaction boundaries. Synthesized in GlobalFoundries 12 nm technology, CLT reduces area by 36.7% relative to PLT while preserving worst-case detection bounds at a minimal detection latency overhead, whereas ILT achieves an 89.2% area reduction suitable for tightly constrained deployments at the cost of a 3.7x higher median detection latency with coarser fault localization. Fault injection campaigns on a RISC-V SoC across 1.2 million scenarios confirm that no fault manifesting as an AXI4 protocol or liveness violation escaped detection, with observed detection latencies consistently bounded by theoretical worst-case predictions.