arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00242cs.CVcs.AIcs.CLcs.RO

CoLT-Drive:用于驾驶 affordance 预测的反事实长尾基准测试与知识保留适应

CoLT-Drive: Counterfactual Long-Tail Benchmarking and Knowledge-Preserving Adaptation for Driving Affordance Prediction

  • NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

Zhengxu Tang, Guofeng Cui, Ziyu Gong, Xiaozhou Zhang, Ruifeng Deng, Chengzhi Qi, Ke Chen, Sachin Patil, Tianjun Xiao, Langechuan Liu, Pichao Wang

AI总结:

本文针对自动驾驶长尾故障的决策级驾驶 affordance 预测问题,提出反事实长尾基准 CoLT-Drive 和知识保留适应框架 KPA,KPA 在 CoLT-Drive 上的动作对准确率达 60.8%,优于基线模型。

AI中文摘要:

自动驾驶中的长尾故障常被归为稀有物体识别错误,本文认为该观点不完整:决策的关键问题不仅是模型是否识别出异常物体,还在于是否能推断该物体如何改变自车可行的高层动作。我们将此问题形式化为决策级驾驶 affordance 预测,即模型将前视图像、自车运动历史和导航指令映射为结构化纵向-横向元动作。为评估该能力,我们提出 CoLT-Drive,这是包含 3536 个样本的反事实长尾基准,它将稀有物体插入原本固定的驾驶场景,以测量模型是否预测出可接受的动作对。为改进可部署的小型视觉语言模型(VLM),我们提出 KPA,即知识保留适应框架,它结合了结构化感知到决策的提示、基于 SLERP 的专家合并,以及 RegMoE(一种感知 regime 的 LoRA 混合专家模块)。KPA 在保留预训练模型开放世界知识的同时,为不同驾驶决策 regime 分配轻量适应容量。在域内驾驶划分和 CoLT-Drive 上的实验显示,KPA 在 CoLT-Drive 上达到 60.8% 的动作对准确率,优于预训练的 Qwen3-VL-2B 基线(50.3%)和 LoRA SFT(32.4%),同时保持有竞争力的域内准确率。我们的基准和代码可在指定 URL 获取。

英文摘要:

Long-tail autonomous driving failures are often framed as rare-object recognition errors. We argue that this view is incomplete: the decision-critical question is not only whether a model recognizes an unusual object, but whether it infers how that object changes the ego vehicle's feasible high-level actions. We formalize this problem as decision-level driving affordance prediction, where a model maps a front-view image, ego-motion history, and navigation command to a structured longitudinal--lateral meta-action. To evaluate this capability, we introduce CoLT-Drive, a 3,536-sample counterfactual long-tail benchmark that inserts rare objects into otherwise fixed driving scenes and measures whether models predict acceptable action pairs. To improve deployable small VLMs, we propose KPA, a knowledge-preserving adaptation framework that combines structured perception-to-decision prompting, SLERP-based expert merging, and RegMoE, a regime-aware LoRA mixture-of-experts module. KPA preserves the pretrained model's open-world knowledge while allocating lightweight adaptation capacity to different driving decision regimes. Experiments on an in-domain driving split and CoLT-Drive show that KPA achieves 60.8\% pair accuracy on CoLT-Drive, outperforming the pretrained Qwen3-VL-2B baseline (50.3\%) and LoRA SFT (32.4\%) while maintaining competitive in-domain accuracy. Our benchmark and code are available at https://huggingface.co/datasets/tangzx2024/CoLT-Drive and https://github.com/tangzhengxu/CoLT-Drive.

补充信息

↑