arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02200cs.CV

RSC-GestureNet:面向中国交通警察手势的可靠性感知选择性因果识别模型

RSC-GestureNet: Reliability-Aware Selective Causal Recognition of Chinese Traffic Police Gestures

  • Academy of Interdisciplinary Studies, The Hong Kong University of Science and Technology(香港科技大学跨学科研究学院)
  • Faculty of Innovation Engineering, Macau University of Science and Technology(澳门科技大学创新工程学院)

机构由 AI 辅助整理,请以论文原文为准。

Cheng Li, Renjun Gao, Boyi Fu

AI总结:

本研究提出RSC-GestureNet,通过对不可靠姿态关节降权、因果聚合时序证据等方式提升中国交通警察手势识别的早期性、稳定性与鲁棒性,在CTPGesture数据集上取得最优性能。

AI中文摘要:

交通警察手势是自动驾驶的安全关键感知线索,可部署的识别器必须从连续全帧视频中因果推理指令,在手臂过渡动作附近保持稳定,且避免过度信任受损的姿态测量。本研究提出RSC-GestureNet,一种面向中国交通警察手势的可靠性感知选择性因果识别模型,该模型将姿态置信度视为一级信号:在图推理过程中对不可靠关节进行降权,因果聚合时序证据,并通过可靠性感知推理规则选择性输出校准后的预测结果。我们进一步引入CTPGesture-C,一个包含7种姿态/RGB退化类型的可复现特征级损坏基准,以及一个RGB级诊断基准,其中损坏帧在识别前由MediaPipe重新处理。在完整的官方CTPGesture v1划分(134424个标注帧和33451个因果窗口)上,RSC-GestureNet实现了93.33±0.24%的准确率、91.71±0.27%的宏F1值、91.69±0.29%的在线宏F1值、98.80±0.07%的Early@10指标、0.153±0.013秒的TTC,且在所有评估方法中具有最优的鲁棒宏F1值。在相同划分和因果协议下,它比复现的交通专用MD-GCN和HLP-GCN基准的宏F1值高3.23-4.11个百分点,在线F1值高2.15-3.07个百分点。这些结果结合校准、选择性风险、统计、自适应分支和图像级重提取分析表明,显式姿态可靠性建模可提升早期、稳定且鲁棒的交通指令识别性能。

英文摘要:

Traffic police gestures are safety-critical perception cues for autonomous driving. A deployable recognizer must infer commands causally from continuous full-frame video, remain stable around transitional arm motion, and avoid over-trusting corrupted pose measurements. This study presents RSC-GestureNet, a reliability-aware selective causal recognizer, for Chinese traffic police gestures. The model treats pose confidence as a first-class signal: unreliable joints are down weighted during graph reasoning, temporal evidence is aggregated causally, and calibrated predictions are selectively emitted through a reliability-aware inference rule. We further introduce CTPGesture-C, a reproducible feature-level corruption benchmark with seven pose/RGB degradation families, and an RGB-level diagnostic in which corrupted frames are reprocessed by MediaPipe before recognition. On the complete official CTPGesture v1 split (134,424 labeled frames and 33,451 causal windows), RSC-GestureNet achieves 93.33+-0.24% accuracy, 91.71+-0.27% macro-F1, 91.69+-0.29% online macro-F1, 98.80+-0.07% Early@10, 0.153+-0.013 s TTC, and the best robust macro-F1 among evaluated methods. Under the same split and causal protocol, it exceeds reproduced traffic-specific MD-GCN and HLP-GCN baselines by 3.23-4.11 macro-F1 points and 2.15-3.07 online-F1 points. These results, together with calibration, selective-risk, statistical, adaptive-branching, and image-level re-extraction analyses, indicate that explicit pose-reliability modeling improves early, stable, and robust traffic-command recognition.

补充信息

↑