arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MoRAL:面向边缘自动驾驶的紧凑型视觉语言模型的传感器接地鸟瞰图推理

MoRAL: Sensor-Grounded BEV Reasoning for Compact VLMs toward Edge-Oriented Autonomous Driving

Ambarish Govindarajulu Kaliamurthi, Kaikai Liu

arXiv 2608.02449首次发表:更新:

AI 中文总结

MoRAL是一种两阶段微调流水线,指导Cosmos-Reason2-2B基于物理编码的BEV表示进行驾驶决策,在参数仅为零样本基准四分之一的情况下,在8类驾驶问题中7类胜出,紧急制动召回率提升至47.8%,可在消费级8GB GPU高效运行。

AI 中文摘要

在资源受限的自动驾驶平台上部署视觉语言模型(VLMs)以进行安全关键的空间推理,需要同时具备紧凑的模型尺寸和可靠的度量接地能力。我们提出了MoRAL(Multimodal Reasoning for Autonomous Language Models,面向自主语言模型的多模态推理),这是一种两阶段微调流水线,用于指导Cosmos-Reason2-2B模型首先读取经物理编码的鸟瞰图(BEV)表示,然后基于该表示进行驾驶决策推理。该BEV图像将激光雷达度量距离编码为色带、将目标类别编码为簇形态、将雷达多普勒速度编码为定向楔形叠加层,将空间感知外化到输入图像中,因此推理时无需使用学习得到的3D骨干网络。第一阶段在60000条接地记录上微调视觉编码器;零样本基准模型无法生成可解析的BEV输出,这证实了其词汇需要明确的训练。第二阶段以Cosmos-Reason2-8B为教师模型生成的57696条思维链记录(涵盖8类驾驶问题)微调完整模型(5200万参数,占总参数的2.4%)。在由Gemma 4(310亿参数)针对人工评审校准的2304个保留的nuScenes帧上进行评估,MoRAL尽管使用的参数仅为零样本8B基准模型的四分之一,却在8类问题中的7类上胜出,在需要结构化多步骤物理推理的问题类型上优势最大。紧急制动召回率从10.8%提升至47.8%,输出退化率从94.1%降至20.8%,且完整流水线无需量化即可在消费级8GB GPU上以42 tok/s的速度运行。这些结果为移动边缘平台上紧凑的、物理接地的VLM推理建立了可复现的基础。

英文摘要

Deploying vision-language models (VLMs) for safety-critical spatial reasoning on resource-constrained autonomous driving platforms requires both compact model size and reliable metric grounding. We present MoRAL (Multimodal Reasoning for Autonomous Language Models), a two-stage fine-tuning pipeline that teaches Cosmos-Reason2-2B to first read a physics-encoded Bird's Eye View (BEV) representation and then reason over it for driving decisions. The BEV image encodes LiDAR metric distance as color bands, object class as cluster morphology, and radar Doppler velocity as directional wedge overlays, externalizing spatial perception into the input image so that no learned 3D backbone is required at inference. Stage 1 fine-tunes the vision encoder on 60,000 grounding records; zero-shot baselines produce no parseable BEV outputs, confirming the vocabulary requires explicit training. Stage 2 fine-tunes the full model (52M parameters, 2.4% of total) on 57,696 chain-of-thought records generated by Cosmos-Reason2-8B as teacher, spanning eight driving question types. On 2,304 held-out nuScenes frames evaluated by Gemma 4 (31B) calibrated against human review, MoRAL wins seven of eight question types over a zero-shot 8B baseline despite using four times fewer parameters, with the largest margins on question types requiring structured multi-step physics reasoning. Emergency braking recall improves from 10.8% to 47.8%, output degeneration falls from 94.1% to 20.8%, and the full pipeline fits a consumer 8 GB GPU at 42 tok/s without quantization. These results establish a reproducible foundation for compact, physics-grounded VLM reasoning on mobile edge platforms.

Comments7 pages, 5 figures, 6 tables. Accepted to the 14th IEEE International Conference on Intelligent Mobile Computing (IEEE IMC 2026), Fukuoka, Japan, July 27-30, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑