arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无需人工标注的X射线安检扫描中任意物体定位

Localize Any Object in X-Ray Security Scans without Human Annotation

Yaqi Cai, Mingxuan Liu, Lorenzo Vaquero, Ning Wang, Nan Pu, Feng Xue, Elisa Ricci, Nicu Sebe

arXiv 2610.07326首次发表:更新:

发表机构

University of Trento; Fondazione Bruno Kessler; Dalian University of Technology; Hefei University of Technology(特伦托大学; 布鲁诺·凯斯勒基金会; 大连理工大学; 合肥工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出LAO-X自监督框架,通过合成图像对和粒度感知监督,无需人工标注即可在X射线安检中实现任意物体定位,显著提升杂乱场景下的定位性能。

AI 中文摘要

在安全关键场所中,X射线安检中的通用物体定位对于自动化威胁检测至关重要。然而,与主导网络规模视觉数据的日常RGB图像不同,X射线扫描呈现出由体积叠加引起的独特颜色模式、模糊边界和组合结构。这些差距阻碍了在Web规模RGB数据上训练的密集感知基础模型直接进行零样本迁移。此外,标注的X射线数据稀缺且需要专家标注,这限制了可泛化的X射线原生模型的训练以及通过微调将RGB基础模型适应于X射线数据的可能性。鉴于这些挑战,由RGB领域数据缩放定律实现的高度泛化感知模型的明亮前景,对于X射线检查来说仍然遥不可及。为此,我们引入了LAO-X,一个自监督适应框架,利用多样化的合成图像-标注对和粒度感知监督来定位X射线扫描中的任意物体。LAO-X首先设计了一个显著性引导的X射线物体挖掘模块来分离不同的物体实例,然后在吸收域中使用物理引导合成。LAO-X进一步结合了遮挡控制课程策略来微调Segment Anything Model 2 (SAM2)定位器,逐步适应具有递增物体数量和重叠级别的X射线扫描。在六个X射线基准上的实验表明,LAO-X显著提高了类别无关定位,在高度杂乱场景中,与SAM2和X射线特定基线相比,实现了2%到23%的mAP提升,完全无需人工标注标签。

英文摘要

Universal object localization in X-ray security inspection is critical for automated threat detection in safety-critical venues. However, unlike everyday RGB images that dominate web-scale visual data, X-ray scans exhibit distinct color patterns, ambiguous boundaries, and compositional structures caused by volumetric superposition. These gaps hinder the direct zero-shot transfer of dense perception foundation models trained on web-scale RGB data. Moreover, annotated X-ray data is scarce and requires expert labeling, limiting both the training of generalizable X-ray native models and the adaptation of RGB foundation models for X-ray data via fine-tuning. Given these challenges, the bright promise of highly generalizable perception models, enabled by data scaling laws in the RGB domain, remains largely out of reach for X-ray inspection. To this end, we introduce LAO-X, a self-supervised adaptation framework that Locates Any Object in X-ray scans using diverse synthesized image--annotation pairs with granularity-aware supervision. LAO-X first designs a saliency-guided X-ray object mining module to separate diverse object instances, which are then used for physics-guided synthesis in the absorbance domain. LAO-X further incorporates an occlusion-controlled curriculum strategy to fine-tune a Segment Anything Model 2 (SAM2) localizer, progressively adapting it to X-ray scans with increasing object counts and overlap levels. Experiments on six X-ray benchmarks show that LAO-X substantially improves category-agnostic localization, achieving 2\% to 23\% mAP gains over SAM2 and X-ray specific baselines in heavily cluttered scenarios, entirely without human-annotated labels.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑