使用细胞自动机架构从兆帧每秒X射线探测器中提取片上布拉格峰
On-chip Bragg peak extraction from MHz frame rate X-ray detectors using a cellular automaton architecture
浏览论文内容
中文总结 AI 辅助
该研究提出一种细胞自动机架构,可在传感器硅片上直接实现X射线布拉格峰的片上定位与补丁提取,能维持高帧频且可准确恢复所有峰,解决了高帧频探测器的数据带宽瓶颈问题。
中文摘要 AI 辅助
高帧频像素探测器可产生的数据量会超出可用的片外带宽,但在许多应用中,仅每帧中的稀疏子集携带相关信息。空间定位事件,包括晶体学中的衍射峰、径迹探测器中的粒子击中、生物成像中的荧光斑点及其他应用,都需要从无特征背景中识别并提取出阈值以上像素的簇。传统上,这在软件算法(如对完整帧进行的连通分量标记)中完成,但在兆帧频下,吞吐量会变得难以承受。我们提出一种轻量级硬件架构,可直接在传感器硅片中执行峰定位和补丁提取,仅传输小像素补丁而非完整帧。在AMD Alveo V80上进行的基于FPGA的测试验证了峰查找模块在实际硬件中的可合成性和时序收敛。该设计用细胞自动机替代全局连通分量标记,采用纯局部、固定迭代次数的邻域操作,消除了使传统方法在片上难以实现的标签存储和等价解析逻辑。我们在远场高能衍射显微镜的X射线布拉格峰检测中验证了该架构,结果显示软件参考找到的每个峰都被成功恢复,且下游重建结果相当。该架构在130nm CMOS工艺中可维持数百千赫兹的帧频,在28nm工艺中则超过1MHz。
英文摘要
High-frame rate pixel detectors can produce data volumes that exceed available off-chip bandwidth, yet in many applications only a sparse subset of each frame carries relevant information. Spatially localized events, including diffraction peaks in crystallography, particle hits in tracking detectors, fluorescence spots in biological imaging, and other applications, all require that clusters of above-threshold pixels be identified and extracted from an otherwise featureless background. Conventionally this is performed in software algorithms such as connected-component labeling on full frames, but at MHz frame rates the resulting throughput becomes prohibitive. We present a lightweight hardware architecture that performs peak localization and patch extraction directly in the sensor silicon, transmitting only small pixel patches rather than complete frames. FPGA-based testing on an AMD Alveo V80 validated the synthesizability and timing closure of the peak-finding module in real hardware. The design replaces global connected-component labeling with a cellular automaton that uses purely local, fixed-iteration neighborhood operations, eliminating the label storage and equivalence-resolution logic that make conventional approaches impractical on-chip. We validate the architecture on X-ray Bragg peak detection for far-field high-energy diffraction microscopy and show that every peak found by a software reference is recovered, with equivalent downstream reconstructions. The architecture sustains several-hundred-kHz frame rates in 130nm CMOS and exceeds 1MHz in 28nm.