AI 中文总结
针对大型二维衍射数据集的数据处理瓶颈,提出将像素分组为可变面积超像素的方法,比较两种方差最小化方法,证明自底向上方法优势,还展示高斯随机投影可加速构建超像素,及超像素用于预处理能大幅加速非负矩阵分解的相位映射。
AI 中文摘要
大型二维衍射数据集通常通过4D-STEM电子衍射或同步加速器X射线束线获取,包含成百上千次测量。机器学习和人工智能在分析这些数据集方面前景广阔,但数据量巨大构成处理瓶颈。裁剪探测器和像素合并是常用的数据缩减方法。本文提出将像素分组并平均为可变面积的超像素,高信息区域密集采样,低信息区域合并为更大超像素,利用数据对称性,适用于二维衍射数据预处理。比较了两种方差最小化方法,证明自底向上的凝聚聚类方法具有更好的扩展性和可解释性。基于距离的方法可通过高斯随机投影加速超像素构建。最后表明在4D-STEM数据集上,超像素用作预处理步骤时,非负矩阵分解的相位映射可加速超100倍。
英文摘要
Large 2D diffraction datasets, consisting of hundreds or thousands of measurements, are commonly acquired with 4D-STEM electron diffraction or at synchrotron X-ray beamlines. Machine learning and artificial intelligence offer great promise for analyzing these datasets. However, the sheer volume of data presents a significant data processing bottleneck. Cropping the detector and pixel binning are standard ways to reduce data size. Here we propose grouping and averaging pixels into superpixels of variable area. High-information regions are sampled densely, while low-information areas are collected into larger superpixels. In the process, symmetries in the data are captured and exploited, making this approach suitable for preprocessing 2D diffraction data. We compare two variance-minimizing methods: K-means clustering (top-down) and agglomerative clustering (bottom-up) and demonstrate superior scaling and interpretability for the bottom-up method. As these methods are distance-based, we demonstrate that the construction of superpixels can be accelerated using Gaussian random projection. Finally we show over 100-fold acceleration for phase mapping with Non-negative matrix factorization on a 4D-STEM dataset when superpixels are used as a preprocessing step.