arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GRU、LSTM与Transformer编码器在自动驾驶系统分类中的敏感性分析

Sensitivity Analysis of GRU, LSTM and Transformer Encoder in Classification of Automated Driving Systems

Bidhya Shrestha, Christos Papadopoulos

arXiv 2607.28665首次发表:更新:

AI 中文总结

本文研究GRU、LSTM、Transformer编码器对Level 2级自动驾驶系统的分类性能,提出含5类损坏的鲁棒性评估框架,发现时间抖动会大幅降低三类模型的宏F1。

AI 中文摘要

自动驾驶系统(ADS)正变得无处不在,未来软件定义车辆(SDV)可能可运行多个ADS,包括原生及售后市场的如该链接的Openpilot。用于独立验证哪个自动驾驶系统处于激活状态的监控系统,对安全监控、合规性监管、保险评估及异常检测至关重要。本文首先评估三种序列分类模型(门控循环单元GRU、长短期记忆LSTM网络、Transformer编码器模型)仅通过车辆远程信息处理数据识别Level 2级自动驾驶系统的有效性,涉及Comma Openpilot、特斯拉Autopilot、凯迪拉克Super Cruise及人工驾驶。所有三种模型在干净数据上均表现优异,干净数据训练时的宏F1分数为GRU 0.92、LSTM 0.90、Transformer编码器模型0.93;威胁匹配训练的宏F1分数为0.904-0.916,仅带来适度的干净数据性能损失。其次,本文提出一种模块化鲁棒性评估框架,通过5类损坏家族及5个严重程度等级(L1-L5)模拟真实远程信息处理数据的退化:连续通道受含累积漂移、跨通道相关噪声及时间抖动的加性高斯白噪声扰动;二元事件信号受突发丢失、过渡延迟、虚假切换及受通信错误启发的跨特征不一致影响。鲁棒性通过宏F1衡量,该指标对各类别权重相等,适用于不平衡多类别评估。评估发现存在明显的故障模式分化:事件级损坏仅轻微降低宏F1(L5时≥0.87),而时间抖动使GRU、LSTM及Transformer编码器模型的宏F1降至0.44-0.50。

英文摘要

Automated driving systems (ADSs) are becoming ubiquitous. Future Software Defined Vehicles (SDVs) may be able to run multiple ADSs, both native and aftermarket such as Comma.ai's Openpilot. Monitoring systems to independently verify which automated driving system is active are important for safety monitoring, regulatory compliance, insurance assessment, and anomaly detection. In this paper, we first evaluate the effectiveness of three sequence-based classification models: Gated Recurrent Units (GRU), Long Short-Term Memory (LSTM) networks, and a Transformer encoder model for identifying Level 2 automated driving systems using vehicle telematics data alone: Comma Openpilot, Tesla Autopilot, and Cadillac Super Cruise, along with manual driving. All three models achieve strong clean-data performance with macro F1-scores of 0.92 (GRU), 0.90 (LSTM), and 0.93 (Transformer encoder model) when trained on clean data; threat-matched training yields 0.904-0.916 macro F1 with only a modest clean-data penalty. Second, we introduce a modular robustness evaluation framework that simulates realistic telematics degradation through five corruption families at five severity levels (L1-L5). Continuous channels are perturbed using additive white Gaussian noise with cumulative drift, correlated cross-channel noise, and temporal jitter. Binary event signals are subjected to burst loss, delayed transitions, spurious toggles and cross-feature inconsistencies inspired by communication errors. Robustness is measured using macro-F1, which gives equal weight to each class and is suitable for imbalanced multiclass evaluation. Our evaluation reveals a sharp failure-mode split: event-level corruptions reduce macro-F1 only slightly (greater than equal to 0.87 at L5), while temporal jitter collapses macro-F1 to 0.44-0.50 across GRU, LSTM, and Transformer encoder model.

Comments7 pages, 6 figures, under review Milcom 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑