TriCCOT: 用于星载空间目标检测的三部分卷积共形Transformer
TriCCOT: Tri-part Convolutional Conformal Transformer for Onboard Space Object Detection
浏览论文内容
中文总结 AI 辅助
TriCCOT提出三部分架构,结合卷积区域提议、共形预测和硬件友好注意力分类器,在DIOR和VDVRaw上实现稳健检测,并成功部署于FPGA。
中文摘要 AI 辅助
地球观测中的星载目标检测受限于有限的计算资源和缺乏完全校正的影像。虽然卷积检测器具有硬件高效性,但它们通常难以从原始且含噪的数据中提取稳健的表征。相反,基于Transformer的模型提供了更强的全局推理能力,但由于二次注意力复杂度和不兼容的操作,它们难以部署在FPGA加速器上。我们提出了TriCCOT,一种用于稳健且可部署的星载目标检测的三部分架构。TriCCOT结合了卷积区域提议网络、共形预测阶段以及Aper-GATES(我们提出的硬件友好的基于注意力的分类器)。区域提议网络生成候选边界框,随后通过共形预测进行扩大,提供无分布的概率覆盖保证。生成的裁剪区域由Aper-GATES处理,该模块通过卷积投影、全局通道统计和硬件友好的门控操作重新表述自注意力,避免了不适合面向CNN加速器的标准Transformer操作。在DIOR和VDVRaw数据集上的实验表明,与FPGA兼容架构相比,TriCCOT实现了有竞争力的检测性能,并增强了对空间模糊和信号相关噪声的稳健性。最后,我们报告了在Xilinx Versal VCK190 FPGA上的完整部署,无需修改底层DPU架构,从而实现了面向星载嵌入式应用的统一CNN-Transformer推理。
英文摘要
Onboard object detection in Earth observation is constrained by limited computational resources and the absence of fully corrected imagery. While convolutional detectors are hardware-efficient, they often struggle to extract robust representations from raw and noisy data. Conversely, transformer-based models provide stronger global reasoning capabilities but remain difficult to deploy on FPGA accelerators due to quadratic attention complexity and non-compatible operations. We introduce TriCCOT, a tri-part architecture for robust and deployable onboard object detection. TriCCOT combines a convolutional region proposal network, a conformal prediction stage, and Aper-GATES, our hardware-friendly attention-based classifier. The region proposal network generates candidate bounding boxes, which are subsequently enlarged via conformal prediction, providing a distribution-free probabilistic coverage guarantee. The resulting crops are processed by Aper-GATES, which reformulates self-attention through convolutional projections, global channel statistics, and hardware-friendly gating operations, avoiding standard transformer operations that are poorly suited to CNN-oriented accelerators. Experiments on the DIOR and VDVRaw datasets demonstrate competitive detection performance and improved robustness to spatial blur and signal-dependent noise when compared to FPGA-compatible architectures. Finally, we report full deployment on a Xilinx Versal VCK190 FPGA without modifying the underlying DPU architecture, enabling unified CNN-Transformer inference for spaceborne embedded applications.
发表机构
- Centre National d’Etudes Spatiales, CNES(法国国家太空研究中心(CNES))
- IRT Saint-Exupéry(IRT圣埃克苏佩里)
机构由 AI 辅助整理,请以论文原文为准。