arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于视觉语言模型的自动驾驶车辆导航GNSS欺骗检测技术发展

Development of Vision-Language Model-based GNSS Spoofing Detection for Autonomous Vehicle Navigation

Mohammed Aldeen, Muhammad Sami Irfan, Sagar Dasgupta, Long Cheng, Mizanur Rahman, Mashrur Chowdhury

arXiv 2607.23962首次发表:更新:

发表机构

Clemson University; The University of Alabama(克莱姆森大学; 阿拉巴马大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究自动驾驶车辆GNSS欺骗检测问题,提出基于视觉语言模型的框架,融合多源数据,经三阶段微调及自适应推理策略,大幅提升检测准确率与效率,并生成数据集验证跨区域泛化能力,为自动驾驶提供实用道路防御层。

AI 中文摘要

自动驾驶车辆依赖全球导航卫星系统(GNSS)进行定位和导航,容易受到欺骗攻击,从而导致车辆转向或不安全操作。本文通过融合前置摄像头视觉数据与车内传感器读数(如速度、加速度、偏航率)和GNSS衍生的操作,开发了首个基于视觉语言模型(VLM)的自动驾驶车辆GNSS欺骗检测框架。我们的方法引入了一个三阶段微调过程,首先对视觉线索进行基础处理,然后在共享语义空间中校准传感器数据,以检测三种攻击场景下预测操作和GNSS衍生操作之间的差异。我们还通过在阿拉巴马州塔斯卡卢萨的公共道路上驾驶一辆装备了时间同步的GNSS、IMU和摄像头日志的仪器车辆,生成了一个独立的真实世界数据集,以验证我们的微调模型在来自训练数据的未见数据上的跨区域泛化能力。在这个数据集上,我们生成了智能欺骗攻击,包括用于错误转向攻击的带有道路网络捕捉的轨迹镜像、用于过冲场景的位置冻结以及用于停车攻击的漂移生成。在这个验证数据集上,零样本VLM基线F1分数在23%到32%之间,而我们的微调模型实现了94%到95%的F1分数。结果表明,我们基于VLM的方法正确分类了每一次错误转向和停车攻击,并在过冲攻击中达到了88%-93%的准确率。此外,我们引入了一种自适应推理策略,将VLM调用减少到14%(计算量减少约86%),并在每4秒窗口内产生65毫秒至73毫秒的响应时间。这些结果表明,使用VLM可以为信号级完整性检查提供一个实用的道路防御层。

英文摘要

Autonomous vehicles (AVs) depend on Global Navigation Satellite Systems (GNSS) for localization and navigation, making them vulnerable to spoofing attacks that can covertly redirect vehicles or induce unsafe maneuvers. In this paper, we develop the first Vision-Language Model (VLM)-based framework for GNSS spoofing detection for autonomous vehicles by fusing front-camera visual data with in-vehicle sensor readings (e.g., speed, acceleration, yaw rate) against GNSS-derived maneuvers. Our approach introduces a three-stage fine-tuning process that first grounds visual cues, and then calibrates sensor data within a shared semantic space to detect discrepancies between predicted and GNSS-derived maneuvers across three attack scenarios. We also generated an independent real-world dataset by driving an instrumented vehicle on public roads in Tuscaloosa, Alabama, equipped with time-synchronized GNSS, IMU, and camera logs to validate cross-regional generalization of our fine-tuned model on unseen data from training data. On this dataset, we then generated intelligent spoofing attacks, including trajectory mirroring with road-network snapping for wrong-turn attacks, position freezing for overshoot scenarios, and drift generation for stop attacks. On this validation dataset, the zero-shot VLMs baseline F1-score ranges from 23% to 32%, whereas our fine-tuned model achieves an F1-score ranging from 94% to 95%. Results show that our VLM-based approach correctly classified every wrong-turn and stop attacks, and attains 88%-93% accuracy for overshoot attacks. Furthermore, we introduce an adaptive inference policy that reduces VLM invocations to 14% (~86% computational reduction) and yields 65ms-73ms per 4s window. These results point to a practical, on-road layer of defense that complements signal-level integrity checks with the use of VLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑