AI 中文总结
本文提出仅用雷达的时序视觉-语言模型Radar4D-VLM,结合提议的时序目标标记等,在K-Radar验证集上表现优异,且雷达标记接口兼容多款冻结语言骨干,为雷达多模态推理奠定基础。
AI 中文摘要
自动驾驶领域的视觉-语言模型主要依赖摄像头和激光雷达,尽管4D雷达对恶劣可见性具有鲁棒性且能直接测量径向速度,却在很大程度上未被作为独立感知模态探索。本文提出Radar4D-VLM,这是一种仅使用雷达的时序视觉-语言模型,无需摄像头或激光雷达输入即可从10个连续的4D雷达点云扫描中进行推理。Radar4D-VLM提取几何上有依据的目标提议,并将雷达证据组织成由目标、场景和运动学标记构成的紧凑层级结构。参数高效的投影器将这些标记映射到冻结的语言骨干网络中,而可审计的预测头则联合建模目标数量、空间分布、运动状态、碰撞风险、语义类别和径向速度。Radar4D-VLM在统一的冻结骨干网络接口内结合了基于提议的时序目标标记、全局场景上下文和显式运动学标记。在序列隔离的K-Radar开发验证集上,其4米处的Top-64提议召回率达到98.13%,分别超过固定格网和均匀随机对照6.40和22.83个百分点。我们在相同的适配预算下,进一步评估了8个冻结骨干网络(包括Qwen、Phi、Mistral、Llama和Gemma)的24次匹配运行。雷达标记接口与所有5个语言模型家族保持兼容,而匹配的对齐、置换和无语言对照显示出传感器依赖性,但对齐的语言监督未带来稳定的直接头增益。这些结果为仅使用雷达的多模态场景和运动推理建立了可复现的基础,同时将接口兼容性与语言监督的益处分离开来。
英文摘要
Vision-language models for autonomous driving primarily rely on cameras and LiDAR, leaving 4D radar largely unexplored as a standalone perceptual modality despite its robustness to adverse visibility and direct measurement of radial velocity. We introduce Radar4D-VLM, a radar-only temporal vision-language model that reasons from ten consecutive 4D-radar point-cloud sweeps without camera or LiDAR input. Radar4D-VLM extracts geometrically grounded object proposals and organizes radar evidence into a compact hierarchy of object, scene, and kinematic tokens. A parameter-efficient projector maps these tokens into frozen language backbones, while auditable prediction heads jointly model object count, spatial distribution, motion state, collision risk, semantic category, and radial velocity. Radar4D-VLM combines proposal-grounded temporal object tokenization, global scene context, and explicit kinematic tokens within a unified frozen-backbone interface. On sequence-isolated K-Radar development validation, its Top-64 proposal recall reaches 98.13% at 4 m, exceeding fixed-lattice and uniform-random controls by 6.40 and 22.83 percentage points, respectively. We further evaluate 24 matched runs spanning eight frozen Qwen, Phi, Mistral, Llama, and Gemma backbones under an identical adaptation budget. The radar-token interface remains compatible across all five language-model families, while matched aligned, permuted, and no-language controls show sensor dependence but no stable direct-head gain from aligned language supervision. These results establish a reproducible foundation for radar-only multimodal scene and motion reasoning while separating interface compatibility from the benefit of language supervision.