arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19088cs.CVcs.AI

通过NMS前预测分布偏移检测目标检测中的后门

Detecting Backdoors in Object Detection via Pre-NMS Prediction Distribution Shift

Longtian Wang, Zhengyu Zhao, Chenhao Lin, Le Yang, Shiwei Wang, Yuhan Zhi, Xiaofei Xie, Chao Shen

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出DistScan框架,通过检测NMS前预测类别分布与训练频率的偏移来检测目标检测模型的后门,在MS-COCO等数据集上较现有方法平均检测准确率提升27.32个百分点。

中文摘要 AI 辅助

部署在安全关键应用中的目标检测模型仍易受到后门攻击,当存在隐藏触发器时会引发定向异常行为。现有检测方法要么依赖触发器反演,要么利用特定于架构的假设,且关键在于,现有代表性方法无法可靠地泛化到场景级攻击,即单个触发器会同时导致场景中所有物体出现异常行为。本文提出DistScan,一种后门检测框架,基于一项简单但此前未被利用的观察:后门注入会系统性地使模型的NMS前预测类别分布偏离其训练类别频率,即便在无任何触发器的干净输入上也是如此。DistScan在干净验证集上聚合中间类别预测结果,若所得分布与训练类别频率显著偏离,则将模型标记为存在后门,无需访问模型权重、无需触发器知识、也无需额外训练。在MS-COCO和PASCAL VOC上针对两种架构及三种场景级攻击场景开展的大量实验表明,DistScan的性能显著优于现有方法,较表现最佳的适用基线提升了27.32个百分点的平均检测准确率。

英文摘要

Object detection models deployed in safety-critical applications remain vulnerable to backdoor attacks that cause targeted misbehaviors when a hidden trigger is present. Existing detection methods either rely on trigger inversion or exploit architecture-specific assumptions, and critically, representative existing methods fail to generalize reliably to scene-level attacks, where a single trigger induces anomalous behavior across all objects in the scene simultaneously. We present DistScan, a backdoor detection framework based on a simple but previously unexploited observation: backdoor injection systematically shifts a model's pre-NMS prediction class distribution away from its training class frequencies, even on clean inputs without any trigger present. DistScan aggregates intermediate class predictions over a clean validation set and flags a model as backdoored if the resulting distribution deviates significantly from the training class frequencies, requiring no model weight access, no trigger knowledge, and no additional training. Extensive experiments on MS-COCO and PASCAL VOC across two architectures and three scene-level attack scenarios demonstrate that DistScan substantially outperforms existing methods, improving average detection accuracy over the best-performing applicable baseline by 27.32 percentage points.

发表机构

  • Singapore Management University(新加坡管理大学)
  • Xi’an Jiaotong University(西安交通大学)

机构由 AI 辅助整理,请以论文原文为准。

↑