发表机构
University of Rzeszów(热舒夫大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出OKS攻击,一种基于决策的黑盒对抗攻击,利用目标关键点相似度作为反馈信号,针对人体姿态估计和动作识别模型,实验显示能有效降低姿态质量并显著降低下游动作识别准确率。
AI 中文摘要
人体姿态估计和基于关键点的动作识别模型越来越多地被部署为视频理解流水线的组成部分,然而它们对对抗攻击的脆弱性仍未得到充分研究。时间上连贯的黑盒攻击先前已在视觉目标跟踪中得到研究,其中攻击反馈可以使用边界框重叠度量(如交并比(IoU))来定义。然而,人体姿态估计产生的是关键点配置而非包围框,这使得基于框的相似性难以适用于度量姿态退化。我们提出了OKS攻击,这是一种基于决策的黑盒攻击,使用目标关键点相似度(OKS)作为攻击反馈信号,直接针对人体姿态的空间结构而非其包围框。在Penn Action数据集上的实验表明,OKS攻击在评估的姿态估计器上持续降低姿态质量,平均OKS下降范围从0.0802到0.1494。在下游跨数据集动作识别评估中,该攻击将准确率降低了6.18到13.86个百分点,并且优于查询匹配的随机噪声扰动。该攻击在自顶向下和单阶段姿态估计模型上均有效。源代码将在该https URL上公开提供。
英文摘要
Human pose estimation and keypoint-based action recognition models are increasingly deployed as components of video understanding pipelines, yet their vulnerability to adversarial attacks remains insufficiently studied. Temporally coherent black-box attacks have been previously studied in visual object tracking, where the attack feedback can be defined using bounding-box overlap measures such as Intersection over Union (IoU). However, human pose estimation produces keypoint configurations rather than enclosing boxes, making box-level similarity poorly suited for measuring pose degradation. We propose OKS Attack, a decision-based black-box attack that uses Object Keypoint Similarity (OKS) as the attack feedback signal, directly targeting the spatial structure of human poses rather than their enclosing boxes. Experiments on the Penn Action dataset show that OKS Attack consistently reduces pose quality across evaluated pose estimators, with mean OKS decreases ranging from 0.0802 to 0.1494. In a downstream cross-dataset action-recognition evaluation, the attack reduces accuracy by 6.18 to 13.86 percentage points and outperforms query-matched random-noise perturbations. The attack is effective across both top-down and single-stage pose estimation models. The source code will be made publicly available at https://github.com/KacperM33/OKS_attack
CommentsAccepted at ACIVS2026