射频成像中基于学习的多对融合的语义重建与三维检测
Semantic Reconstruction and 3-D Detection via Learned Multi-Pair Fusion in RF Imaging
浏览论文内容
中文总结 AI 辅助
该研究针对各向异性多站射频成像问题,提出用三维U-Net融合多对Tx-Rx的重建结果以实现语义重建与3-D检测,新增未知类别提升对新对象的识别能力,其性能优于经典强度重建。
中文摘要 AI 辅助
我们考虑具有各向异性的多站射频成像问题,其中点的反射取决于发射(Tx)和接收(Rx)阵列的位置。目标是通过有限的语义类别对视场的体素进行标记,并将其分组为目标实例。对于每对Tx-Rx的图像形成,我们采用标准的逆问题求解器,并将得到的每对重建结果输入到经过训练的三维(3-D)U-Net中,该网络隐式执行融合并显式执行每体素分类。在受控的欠定多站设置下,我们考虑以下图像形成方法:来自单个确定性快照的反投影(BP)和最小绝对收缩与选择算子(LASSO),以及来自多个衰落快照的非相干BP和组-LASSO。对于每种成像方法,我们训练一个单独的U-Net,该网络融合六对Tx-Rx(其输入通道),并为每个体素分配类别上的概率向量。取概率最高的类别得到标记体积——即语义重建。随后通过几何后处理(聚类和主成分分析)得到目标实例及其定向边界框。在宽范围的信噪比下,语义重建(通过分割交并比与真值评分)和由此产生的3-D检测的退化远缓于经典的强度重建:特别是,检测在强度重建已失效的噪声水平下仍保持可靠。由于真实场景包含网络未训练过的类别对象,我们添加了通过离群值曝光训练的显式未知类别,该类别将保留的新对象标记为未知,而非通过重建形状将其错误标记为已知类别。
英文摘要
We consider a multistatic radio-frequency imaging problem with anisotropy, in which the reflection from a point depends on the positions of the transmit (Tx) and receive (Rx) arrays. The goal is to label the voxels of a field of view by a finite set of semantic classes and to group them into object instances. For the image formation of each Tx--Rx pair we apply a standard inverse-problem solver, and we feed the resulting per-pair reconstructions into a trained three-dimensional (3-D) U-Net that performs the fusion implicitly and the per-voxel classification explicitly. On a controlled, under-determined multistatic setup, we consider the following image formation methods: back-projection (BP) and the least absolute shrinkage and selection operator (LASSO) from a single deterministic snapshot, and incoherent BP and group-LASSO from multiple fading snapshots. For each imaging method we train a separate U-Net that fuses the six Tx--Rx pairs (its input channels) and assigns each voxel a probability vector over the classes. Taking the most probable class gives a labeled volume---the semantic reconstruction. Object instances and their oriented bounding boxes then follow by geometric post-processing (clustering and principal-component analysis). Across a wide range of signal-to-noise ratio, the semantic reconstruction (scored against ground truth by segmentation intersection-over-union) and the resulting 3-D detection degrade far more gracefully than the classical intensity reconstruction: the detection in particular stays reliable well into noise levels at which that reconstruction has dissolved. Because real scenes contain objects of classes the network was not trained on, we add an explicit unknown class trained by outlier exposure, which labels held-out novel objects as unknown instead of mislabeling them as a known class by reconstructed shape.
发表机构
- Technische Universität Berlin(柏林工业大学)
机构由 AI 辅助整理,请以论文原文为准。