arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Dryas:用于高速互连追踪与分析的可重编程引擎

Dryas: A Reprogrammable Engine for High-Speed Interconnect Tracing and Analysis

Manuel Bröchin, Tom Kuchler, Michael Giardino, David Cock, Timothy Roscoe

arXiv 2608.12934首次发表:更新:

AI 中文总结

Dryas是一款开源可重编程高速互连分析工具,基于NFAs与STEs实现,可在不干扰应用的情况下追踪事件,用于互连调试与缓存行为分析,具备高可扩展性。

AI 中文摘要

现代计算系统中异构组件的激增催生了更高带宽、更低延迟的新型互连技术,这些接口与协议极为复杂,开发、调试和分析基于FPGA的实现需要大量工程工作;且功能实现完成后,控制器及相关软件的优化需处理可能达数百GB的追踪数据。本文提出Dryas,一款用于分析此类互连的开源工具,我们以极少硬件资源开发了该工具,同时实现了一款30 GiB/s、延迟200 ns的超高速低延迟互连的FPGA实现。借助运行时可重编程的覆盖引擎,我们可在互连满负荷运行时检查其以发现罕见、复杂或瞬态事件;该过滤引擎基于非确定性有限自动机(NFAs),通过状态转换元件(STEs)高效实现,支持以缓存行粒度追踪事件,且可在不到1秒内更改过滤器,无需重新编程FPGA或干扰运行中的应用。这些数据不仅可用于调试互连自身的实现,还可分析加速应用的行为。我们研究了使用NFAs的数学基础,描述其在真实一致CPU-FPGA研究平台上的实现,随后评估Dryas对不同规模NFAs的可扩展性,并给出两个不同用例:调试互连的FPGA实现与分析缓存行为。

英文摘要

The proliferation of heterogeneous components in modern computing systems has been accompanied by new higher bandwidth and lower latency interconnects. These interfaces and protocols are enormously complex and the process of developing, debugging, and analyzing FPGA-based implementations requires significant engineering work. Moreover, once a functional implementation is completed, optimization of the controller and associated software requires processing potentially hundreds of gigabytes of trace data. In this paper, we present Dryas, an open source tool for analyzing such an interconnect. We developed our tool, using minimal hardware resources, alongside an FPGA implementation of a very high speed, low latency (30~GiB/s, 200~ns) interconnect. With our run-time reprogrammable overlay engine we can inspect this interconnect to find rare, complex, or transient events even at full operation. This filtering engine is based on non-deterministic finite automata (NFAs), efficiently implemented using state transition elements (STEs), allowing us to trace events at a cache-line granularity. Moreover we can change the filters in less than a second, without reprogramming the FPGA or interfering with the running application. This data enables not only debugging the implementation of the interconnect itself, but analyzing the behavior of accelerated applications. We examine the mathematical basis for using NFAs and describe their implementation on a real coherent CPU-FPGA research platform. We then evaluate the scalability of Dryas for various size NFAs, followed by two different use cases: debugging FPGA implementation of the interconnect and analyzing cache behavior.

Comments12 pages, 12 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑