arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DH-VLM:面向自动驾驶的双时域协同隐层推理框架

DH-VLM: Dual-Horizon Cooperative Latent Reasoning for Autonomous Driving

Ziyi Song, Chen Xia, Hang Yu, Sheng Zhou, Zhisheng Niu

arXiv 2608.09333首次发表:更新:

发表机构

Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出DH-VLM双时域协同隐层推理框架,结合基础设施与自车实现非对称语义协同,构建协同QA数据集支撑推理,在规划性能、通信成本等指标上优于现有方法,为协同自动驾驶提供实用鲁棒范式。

AI 中文摘要

面向自动驾驶的大规模语言模型可提升全局理解能力与长时域规划能力,但部署在孤立车辆上时,有限的感知范围与遮挡问题会限制可靠决策,且可观的计算与延迟开销导致车载部署不切实际。协同驾驶通过利用外部智能体进行信息交换提供了潜在解决方案,但现有方法在实际约束下的语义推理能力仍有限。为应对这些挑战,本文提出DH-VLM,一种双时域协同隐层推理框架,可实现基础设施与自车之间的非对称语义协同。基础设施聚合多层隐状态以形成全局推理时域的隐层引导,该引导通过基础设施驱动的隐层演化机制集成到自车模型中,用于条件隐层优化,使自车能利用长程上下文理解,同时在其本地规划时域内保留自主决策能力。此外,本文构建了面向协同的问答(QA)数据集,涵盖基础场景理解与自车个性化理解,以支持反事实与安全感知推理。大量实验表明,DH-VLM实现了最优规划性能,在L2误差上较之前最优方法提升14.6%,碰撞率降低26.9%;与基于查询的端到端协同驾驶方法相比,本文方法降低57.3%的通信成本与25.5%的GPU内存使用,同时保持对基础设施引导错误的强鲁棒性,为协同自动驾驶提供了实用且鲁棒的范式。

英文摘要

Large-scale language models for autonomous driving enable enhanced global understanding and long-horizon planning. However, when deployed in isolated vehicles, limited sensing range and occlusions restrict reliable decision-making, and the substantial computational and latency overhead makes on-board deployment impractical. Cooperative driving provides a potential solution by leveraging external agents for information exchange, but existing methods remain limited in semantic reasoning capability under practical constraints. To address these challenges, we propose DH-VLM, a dual-horizon cooperative latent reasoning framework that enables asymmetric semantic cooperation between the infrastructure and ego vehicle. The infrastructure aggregates multi-layer hidden states to form a global-reasoning horizon latent guidance, which is integrated into the ego model through an Infrastructure-Driven Latent Evolution mechanism for conditional latent refinement. This enables the ego vehicle to leverage long-range contextual understanding while preserving autonomous decision-making within its local planning horizon. Furthermore, we construct a cooperation-oriented question-answer (QA) dataset covering fundamental scene understanding and ego-personalized comprehension to support counterfactual and safety-aware reasoning. Extensive experiments demonstrate that DH-VLM achieves state-of-the-art planning performance, outperforming the previous state of the art by 14.6% in L2 error and 26.9% in collision rate. Compared with query-based end-to-end cooperative driving methods, our approach reduces the communication cost by 57.3% and GPU memory usage by 25.5%, while maintaining strong robustness against infrastructure guidance errors, providing a practical and robust paradigm for cooperative autonomous driving.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑