发表机构
University of Michigan(密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出STAR-VLM框架,利用汽车雷达监督提升时空视觉语言模型的运动推理与度量速度估计能力,在驾驶场景相关任务上实现了优于特定任务方法的最先进性能。
AI 中文摘要
视觉语言模型(VLMs)正成为具身智能的核心组成部分,在自动标注和端到端自动驾驶中的应用日益广泛。然而,现有提升VLMs时空推理能力的方法往往依赖复杂的预处理流程、昂贵的人工标注或合成数据,这限制了可扩展性并引入了潜在的仿真到现实的差距。此外,尽管这些方法已改善了时空理解,但仍缺乏针对动态场景的强度量推理能力,例如以真实世界单位估计物体运动。先前研究探索了基于LiDAR的度量深度监督以增强空间感知,但未直接解决时间推理问题。我们提出STAR-VLM,一种汽车雷达监督框架,通过运动推理和度量速度估计增强用于自动驾驶的时空VLMs。汽车雷达是一种低成本且广泛部署的传感器,能通过距离和多普勒测量提供互补的时空监督。在训练期间将这些测量作为无标签的真实值,STAR-VLM可提升VLMs的度量时空推理能力。通过在驾驶场景上的实验,我们表明STAR-VLM在运动分类和度量速度估计上均达到了最先进的性能,甚至优于为每个任务设计的特定任务方法。这些结果凸显汽车雷达是一种可扩展且具成本效益的监督源,用于构建面向现实世界自动驾驶的感知度量的时空VLMs。
英文摘要
Vision-language models (VLMs) are emerging as a key component of embodied intelligence, with growing applications in auto-labeling and end-to-end autonomous driving. However, existing approaches for improving spatiotemporal reasoning in VLMs often rely on complex preprocessing pipelines, expensive human annotations, or synthetic data, which limit scalability and introduce potential sim-to-real gaps. Moreover, although these methods have improved spatiotemporal understanding, they still lack strong metric reasoning capabilities for dynamic scenes, such as estimating object motion in real-world units. Prior work has explored LiDAR-based metric depth supervision to enhance spatial perception, but it does not directly address temporal reasoning. We introduce STAR-VLM, an automotive radar-supervised framework that enhances spatiotemporal VLMs with motion reasoning and metric velocity estimation for autonomous driving. Automotive radar is a low-cost and widely deployed sensor that provides complementary spatiotemporal supervision through range and Doppler measurements. By leveraging these measurements as label-free ground truth during training, STAR-VLM improves the metric spatiotemporal reasoning ability of VLMs. Through experiments on driving scenarios, we show that STAR-VLM achieves state-of-the-art performance on both motion classification and metric velocity estimation, outperforming even task-specific methods designed for each task. These results highlight automotive radar as a scalable and cost-effective source of supervision for building metric-aware spatiotemporal VLMs for real-world autonomous driving.