arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31217cs.NI

深度强化学习用于部分可观测V2X数据下的异常行为检测

Deep Reinforcement Learning for Misbehavior Detection Under Partially Observable V2X Data

  • Centre Tecnològic de Telecomunicacions de Catalunya (CTTC/CERCA)(加泰罗尼亚电信技术中心)

机构由 AI 辅助整理,请以论文原文为准。

Roshan Sedar, Charalampos Kalalas

AI总结:

本文针对部分可观测V2X数据下的异常行为检测,提出基于深度强化学习的自适应检测框架,在VeReMi数据集上优于XGBoost基线,但易受利用数据缺失的逃避攻击影响。

AI中文摘要:

车联网(V2X)系统中的异常行为检测对于确保交换消息的语义正确性以及防止虚假信息的传播至关重要。现有的以数据为中心的异常行为检测方法在很大程度上依赖于统计验证或监督式机器学习模型,这些方法隐含地假设V2X数据流是完全可观测的。然而,在实践中,由于硬件故障、间歇性连接和环境遮挡,车辆环境本质上是部分可观测的。此外,数据缺失本身可能被对手策略性地利用以逃避检测。在本文中,我们研究了在不完整V2X观测下的异常行为检测,并提出了一种基于深度强化学习(DRL)的检测框架,该框架能够利用不完整数据学习自适应策略。我们进一步引入了一种对抗性威胁模型,在该模型中,攻击者利用或故意诱导数据缺失以逃避检测,包括通过自然遮挡进行逃避以及对抗性特征抑制。在VeReMi数据集上针对各种缺失模式进行的大量实验表明,在自然部分可观测性条件下,DRL显著优于强大的XGBoost基线。然而,结果也揭示了一个关键漏洞:DRL策略可能极易受到策略性利用自然缺失的逃避攻击的影响。相比之下,在直接特征抑制下,DRL相比静态基于树的模型表现出更渐进的性能退化。

英文摘要:

Misbehavior detection in vehicle-to-everything (V2X) systems is essential for ensuring the semantic correctness of exchanged messages and preventing the dissemination of falsified information. Existing data-centric misbehavior detection approaches largely rely on statistical validation or supervised machine learning models under the implicit assumption of fully observable V2X streams. In practice, however, vehicular environments are inherently partially observable due to hardware failures, intermittent connectivity, and environmental occlusions. Moreover, missingness itself can be strategically exploited by adversaries to evade detection. In this paper, we study misbehavior detection under incomplete V2X observations and propose a deep reinforcement learning (DRL)-based detection framework that learns adaptive policies with incomplete data. We further introduce an adversarial threat model in which attackers exploit or deliberately induce missingness to evade detection, including evasion via natural occlusions and adversarial feature suppression. Extensive experiments conducted on the VeReMi dataset under various missingness patterns demonstrate that DRL significantly outperforms a powerful XGBoost baseline under natural partial observability. However, results also reveal a critical vulnerability: DRL policies can be highly susceptible to evasion attacks that strategically exploit natural missingness. In contrast, DRL exhibits more gradual degradation under direct feature suppression compared to static tree-based models.

↑