arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02806cs.CRcs.CV

面向安全关键型基于视频的感知系统的快速物体移除攻击

Fast Object Removal Attacks on Safety-Critical Video-based Perception Systems

Mohammad Imtiaz Hasan, M Sabbir Salek, Nathan Jones, Mashrur Chowdhury, Rong Ge

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种端到端快速针对性物体移除攻击框架,可近实时破坏基于视频的安全关键感知系统,在SC-CVT实验中攻击成功率达94.48%,难被篡改检测模型识别,为防御此类攻击提供依据。

中文摘要 AI 辅助

基于视频的感知系统所提供的数据,智能交通系统(ITS)支撑着提升道路安全的安全关键型应用。然而,攻击者可能会操纵视频帧,以破坏下游感知模块,进而导致安全关键型功能失效,增加弱势道路使用者的风险。本文提出了一种新颖的攻击模型以及用于对基于视频的安全关键型系统实施近实时针对性物体移除攻击的端到端框架。该端到端攻击流程包含四个阶段:在每一帧中定位目标、从先前帧中检索连贯的图像块、利用上下文感知的alpha合成法将这些图像块融合,以及重建被攻击的帧。在南卡罗来纳州车联网测试平台(SC-CVT)的一个交叉路口开展的实验表明,重建后的帧与原始帧具有很高的全局相似度,帧级峰值信噪比(PSNR)高于40分贝,结构相似性指数测度(SSIM)高于0.996。采用基于YOLO的检测器时,该攻击可将物体检测率降低多达97.59%,并实现94.48%的帧级攻击成功率。在评估所用的检测器和帧分辨率下,在GPU硬件上的平均执行时间为每帧0.074至0.172秒,显示出测试阶段的近实时性能。使用多个预训练的篡改检测模型开展的取证评估显示,这些模型区分重建帧与真实帧的能力有限。研究结果表明,基于视频的感知系统易受隐蔽的物体移除攻击,此类攻击会通过降低物体可检测性来削弱安全关键型应用的性能。这些研究结果有助于制定针对威胁安全关键型应用的对抗性物体移除攻击的缓解策略,例如基于视觉的行人安全系统。

英文摘要

By leveraging data from video-based perception systems, intelligent transportation systems (ITS) support safety-critical applications that improve road safety. However, adversaries may manipulate video frames to compromise downstream perception modules, causing failures in safety-critical functions and increasing risks to vulnerable road users. This paper presents a novel attack model and an end-to-end framework for near-real-time targeted object removal attack on a video-based safety-critical system. The end-to-end attack pipeline consists of four stages: localizing targets in each frame, retrieving coherent patches from earlier frames, blending them using context-aware alpha compositing, and reconstructing attacked frames. Experiments at an intersection on the South Carolina Connected Vehicle Testbed (SC-CVT) show that reconstructed frames have high global similarity to the originals, with frame-level Peak Signal to Noise Ratio (PSNR) above 40 dB and Structural Similarity Index Measure (SSIM) above 0.996. Using the YOLO-based detector, the attack reduces object detections by up to 97.59% and achieves a frame-level attack success rate of 94.48%. Across the evaluated detectors and frame resolutions, the mean execution time ranges from 0.074 to 0.172 seconds per frame on GPU hardware, indicating near-real-time performance in testing. The forensic evaluation using several pretrained tamper-detection models shows limited ability to distinguish reconstructed from authentic frames. The findings suggest that video-based perception is vulnerable to stealthy object removal attacks that can degrade the performance of safety-critical applications by reducing object detectability. These findings can help develop mitigation strategies against adversarial object removal attacks that threaten safety-critical applications, such as vision-based pedestrian safety systems.

发表机构

  • Clemson University(克莱姆森大学)
  • National Center for Transportation Cybersecurity and Resiliency(国家交通网络安全与韧性中心)

机构由 AI 辅助整理,请以论文原文为准。

↑