AI 中文总结
研究针对小芯片仿真中链路建模简化问题,提出DICE方法,在gem5中进行运行时PHY建模,能捕获端到端芯片间数据路径,解决链路动态行为被忽略问题,提升小芯片仿真准确性。
AI 中文摘要
单片多核扩展越来越受到功率/热限制、良率以及制造和测试成本上升的制约。小芯片设计通过将大芯片划分为更小的部分(通常是多个核心复合体芯片和一个I/O芯片),通过高带宽物理结构(PHY)连接来应对这些挑战。随着带宽和布线密度的扩展,短距离链路接近信号完整性极限,需要更强的链路级可靠性机制。然而,现有最先进的仿真基础设施常使用过于简化的固定延迟模型来近似芯片间链路。这种抽象忽略了PHY固有的动态、运行时依赖行为,导致芯片间数据包级定时和IPC等高级性能指标失真。我们展示了DICE,一种在gem5中进行仿真时的运行时PHY建模,可捕获端到端芯片间数据路径。
英文摘要
Scaling monolithic multicores is increasingly constrained by power/thermal limits, yield, and rising manufacturing and testing costs. Chiplet designs address these challenges by partitioning large dies into smaller parts (typically multiple core-complex dies and an I/O die) linked via high-bandwidth physical fabrics (PHY). As bandwidth and wiring density scale, however, these short-reach links are pushed closer to their signal-integrity limits, increasing susceptibility to noise, crosstalk, and channel loss, motivating stronger link-level reliability mechanisms such as forward error correction (FEC). Despite this trend, state-of-the-art simulation infrastructures often approximate inter-chiplet links using oversimplified, fixed-latency models. Such abstractions overlook the inherently dynamic, runtime-dependent behavior of the PHY -- including channel conditions (e.g., signal-to-noise ratio shifts, signal crosstalk, clock jitter), iterative decoder convergence and packet retransmissions, and application dynamics (e.g., LLC-misses that travel across chiplet boundaries) -- all of which are hard to determine offline. We show that neglecting these effects distorts inter-chiplet packet-level timing and high-level performance metrics such as IPC, leading to off-trend simulation results. We present DICE, an in-simulation, runtime PHY modeling in gem5 that captures the end-to-end inter-chiplet datapath, including QC-LDPC encoding/decoding, PAM4 modulation, lossy-channel transmission, LLR-based demodulation, adaptive packet re-sending, and PHY-level flow control between chiplets.