arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22078cs.CV

CommandLM:用于自动驾驶车辆的数据驱动行为级描述符

CommandLM: Data driven behavior level descriptor for ego vehicles

Boris Tokic, Constantin Selzer, Fabian B. Flohr

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对自动驾驶系统中可解释行为级决策需求,提出多模态大语言模型CommandLM,通过融合多传感器数据生成行为描述,经特定训练和适配器处理,实验表现优于基线,人工评估显示其输出可助力行为审计,实现高效透明的行为级理解。

中文摘要 AI 辅助

随着自动驾驶系统向实际部署迈进,可解释的行为级决策对于安全、信任和监管至关重要。我们引入了CommandLM,这是一个多模态大语言模型,它从融合的多传感器数据中为自动驾驶车辆生成简洁、人类可读的行为描述。我们的模型通过连接到量化的、LoRA微调的大语言模型的Q-Former适配器来处理来自激光雷达和多摄像头输入的时间融合鸟瞰图表示。在CommandLM-nuScenes数据集上训练后,CommandLM生成适用于规划器监督和安全审计的意图感知、可解释的字幕。实验证明了强大的语言和行为对齐,在CIDEr指标上达到0.67,BERT-F1指标上达到0.88,大幅超越BLIP-2基线。在人工评估中,58%的生成描述被评为准确、高效和符合规则。CommandLM可解释的输出使下游验证系统能够识别并纠正此类情况,是透明行为审计的有效工具。这些结果表明,将多模态融合与语言推理相结合可为自动驾驶带来高效且透明的行为级理解。

英文摘要

As autonomous driving systems move toward real-world deployment, interpretable, behavior-level decision-making is essential for safety, trust, and regulation. We introduce CommandLM, a multimodal large language model that generates concise, human-readable behavior descriptions for ego vehicles from fused multi-sensor data. Our model processes temporally fused bird's-eye view representations from LiDAR and multi-camera inputs via a Q-Former adapter connected to a quantized, LoRA-fine-tuned large language model. Trained on our CommandLM-nuScenes dataset, CommandLM produces intent-aware, interpretable captions suitable for planner supervision and safety auditing. Experiments demonstrate strong linguistic and behavioral alignment, achieving CIDEr 0.67, and BERT-F1 0.88, substantially outperforming the BLIP-2 baseline (CIDEr 0.52, BERT-F1 0.86). In human evaluation, 58% of the generated descriptions were rated accurate, efficient and rule-compliant, confirming their real-world plausibility. While the remaining descriptions may not always select the most efficient, goal-oriented behavior, CommandLM's interpretable outputs enable downstream validation systems to identify and correct such cases, making it an effective tool for transparent behavior auditing. These results show that integrating multimodal fusion with language reasoning yields efficient and transparent behavior-level understanding for autonomous driving. We release our code and dataset at: https://github.com/b-tok/CommandLM

发表机构

  • Munich University of Applied Sciences(慕尼黑应用技术大学)

机构由 AI 辅助整理,请以论文原文为准。

↑