arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

驱动思考:VLA推理-轨迹一致性的运行时监测

Drive the Thoughts: Runtime Monitoring of VLA Reasoning-Trajectory Consistency

Tian Yu, Lu Feng, Sebastian Elbaum

arXiv 2608.29583首次发表:更新:

AI 中文总结

本文研究VLA的CoT能否支持AV的运行时监测,构建了DriveAlignBench数据集,提出带GPT-5.5的车道相关F-LLM监测器,F1达0.75并发布相关资源。

AI 中文摘要

自动驾驶车辆(AV)在复杂环境中运行,故障会产生严重后果。感知与规划领域的复杂机器学习模型是应对部分复杂性的关键,但它们的黑盒特性给验证与确认(V&V)带来了挑战。近期视觉-语言-动作(VLA)模型在AV中的集成提供了独特机遇:这类模型除生成轨迹外,还会产生显式的思维链(CoT)以解释其底层逻辑,该CoT为交叉校验模型输出、检测可能暴露不安全或非预期行为的不一致性提供了丰富依据。本文评估近期开源驾驶VLA的CoT是否可支持此类监测。我们构建了DriveAlignBench,这是一个来自NVIDIA的Alpamayo 1.5 VLA的AV专用数据集,包含150个CoT-轨迹对,我们对其进行了可靠性、轨迹一致性及安全性的手动标注。分析显示,33.3%的CoT不可靠;在可靠的CoT中,生成的轨迹与CoT一致的情况占74%。利用该潜力,我们提出将CoT-轨迹一致性检查集成到运行时监测器中。该检查并非易事:CoT表达开放词汇、与场景相关的驾驶承诺,而轨迹是低级自我运动序列,其语义依赖于道路几何与运动上下文。为弥合这一差距,我们开发了一系列自动一致性监测器。我们的最优监测器为带GPT-5.5的车道相关F-LLM,其F1值达0.75,较最强的原始路点LLM基线提升了0.13的绝对F1值,较基于规则的监测器提升了0.38。我们在该httpsURL发布了DriveAlignBench、监测器实现及标注工具。

英文摘要

Autonomous vehicles (AVs) operate in complex environments where failures are consequential. Sophisticated machine learning models for perception and planning are key to overcoming at least part of that complexity, but their black-box nature complicates validation and verification (V&V). The recent integration of Vision-Language-Action (VLA) models into AVs introduces a unique opportunity: besides generating trajectories, these models produce an explicit Chain-of-Thought (CoT) explaining their underlying rationale. This CoT provides a rich specification to cross-check model outputs and detect inconsistencies that may expose unsafe or unintended behavior. This paper assesses whether CoTs from a recent open driving VLA can support such monitoring. We curate DriveAlignBench, a specialized dataset from NVIDIA's Alpamayo 1.5 VLA for AVs containing 150 CoT-trajectory pairs, which we manually annotate for reliability, trajectory consistency, and safety. Our analysis reveals that 33.3% of CoTs are unreliable. Among reliable CoTs, the generated trajectory is consistent with the CoT in 74% of cases. Leveraging this potential, we propose integrating a CoT-trajectory consistency check into a runtime monitor. The check is nontrivial: CoTs express open-vocabulary, scene-relative driving commitments, while trajectories are low-level ego-motion sequences whose semantics depend on road geometry and motion context. To bridge this gap, we develop a family of automated consistency monitors. Our best monitor, lane-relative F-LLM with GPT-5.5, achieves F1 = 0.75, improving over the strongest raw-waypoint LLM baseline by +0.13 absolute F1 and over a rule-based monitor by +0.38. We release DriveAlignBench, the monitor implementations, and annotation tools at https://github.com/776styjsu/drive-the-thoughts.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑