doPlan:面向自动驾驶多阶段语言条件规划的变视界数据集
doPlan: A Variable-Horizon Dataset for Multi-Stage Language-Conditioned Planning in Autonomous Driving
浏览论文内容
中文总结 AI 辅助
提出首个公开人工标注的真实世界数据集doPlan,基于nuPlan构建,含5154条乘客指令,覆盖多阶段、事件条件等长视界意图,评估四个模型发现其对语言敏感但行为不一致,强调需将持久意图与连续规划连接。
中文摘要 AI 辅助
通过自然语言与乘客交互的自动驾驶车辆必须超越即时命令进行推理。乘客意图可能涵盖多个行为阶段,依赖于未来事件,涉及周围智能体或地标,并在驾驶条件演变时保持相关性。现有的语言驱动数据集大多侧重于短时、局部的交互,使得这些较长视界的乘客意图形式相对未被充分探索。我们引入了doPlan,据我们所知,这是首个公开可用的、人工标注的真实世界数据集,旨在研究乘客语言作为持久任务上下文。基于nuPlan构建,doPlan包含5,154条人工撰写的乘客指令,覆盖169.1小时的累计指令对齐上下文,对应50.9小时的独特驾驶,标注窗口范围为30.0至508.8秒。这些标注捕捉了即时、延迟、事件条件、持久和多阶段的乘客意图。该数据集、标注界面及支持资源可在以下网址公开获取:此https URL。我们评估了四个语言条件驾驶模型,发现对乘客语言的敏感性并不能可靠地转化为与请求方向一致的行为。更广泛地说,在2,161个具有匹配未来机动操作的示例中,第一个相关机动操作在评估点后的中位时间为24.6秒,且仅有9.8%发生在模型常见的5秒预测视界内。这些发现强调了将持久乘客意图与连续规划决策连接起来的必要性。doPlan为研究未解决目标如何被保留、如何基于演变场景进行落地以及如何跨多个阶段进行跟踪(包括规划器如何确定未来目标何时与当前计划相关)提供了一个环境。
英文摘要
Autonomous vehicles interacting with passengers through natural language must reason beyond immediate commands. Passenger intent may span multiple stages of behavior, depend on future events, refer to surrounding agents or landmarks, and remain relevant as driving conditions evolve. Existing language-enabled driving datasets largely focus on short, localized interactions, leaving these longer-horizon forms of passenger intent comparatively underexplored. We introduce doPlan, to our knowledge the first publicly available, human-annotated real-world dataset designed to study passenger language as persistent task context. Built on nuPlan, doPlan contains 5,154 human-written passenger instructions spanning 169.1 hours of cumulative instruction-aligned context over 50.9 hours of unique driving, with annotation windows ranging from 30.0 to 508.8 s. The annotations capture immediate, deferred, event-conditioned, persistent, and multi-stage passenger intent. The dataset, annotation interface, and supporting resources are publicly available at https://github.com/Mi3-Lab/doPlan. We evaluate four language-conditioned driving models and find that sensitivity to passenger language does not reliably translate into behavior consistent with the requested direction. More broadly, among 2,161 examples with a matched future maneuver, the first associated maneuver occurs a median of 24.6 s after the evaluation point, and only 9.8% occur within the models' common 5 s prediction horizon. These findings highlight the need to connect persistent passenger intent with successive planning decisions. doPlan provides a setting for studying how unresolved goals can be retained, grounded in evolving scenes, and tracked across multiple stages, including how a planner determines when a future goal becomes relevant to the current plan.
发表机构
- Laboratory for Intelligent & Safe Automobiles (LISA), University of California, San Diego(加州大学圣迭戈分校智能与安全汽车实验室(LISA))
- Machine Intelligence, Interaction, and Imagination (Mi) Laboratory, University of California, Merced(加州大学默塞德分校机器智能、交互与想象实验室(Mi))
机构由 AI 辅助整理,请以论文原文为准。