arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

X-Planner:具身智能的事件结构化任务规划

X-Planner: Event-Structured Task Planning for Embodied Intelligence

Howard Lu, Shalfun Li, Porter Pan, Cris, Lumen, Cyril, Eric Hu, Lily Li, Maeve Zhang, Rain Sun, Robert Wang, KZ Zheng, Viggo Chen, Tim Ding, Regsis Cheng, YJ Xiao, Kian, Hai Lin, Alan Song, Elise Ma, Gody Li, Victor Yao, Yohann Tang, Ingrid Yu, Jason He, James Wang, Ryan Yu, Ping Yang, Chris Pan, Vincent Chen, Roy Gan, Hao Wang, Qian Wang

arXiv 2609.25187首次发表:更新:

发表机构

X Square Robot Team(X Square Robot 团队)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

X-Planner提出事件结构化任务规划前端,结合离散与潜在接口及阶梯式解码,改善具身智能的推理监督与表示,在离线评估中排名第二并优于真实机器人基线。

AI 中文摘要

任务规划在长时程操作中连接高层指令与可执行行为,然而现代的视觉-语言-动作(VLA)系统往往使这一中间结构隐式化。现有的思维链(CoT)规划器也倾向于依赖粗略的任务级注释,或逐词元地序列化冗长的推理轨迹。我们提出X-Planner,一个规划前端,旨在同时解决具身推理的监督与表示问题。我们的规划数据在层级粒度下结合了Ego、UMI和遥操作,并具有依赖数据源的注释深度。接管时注释和人为设计的失败用于监督持续的错误识别。在模型方面,共享的VLM主干暴露了两种事件结构化规划形式:一种离散接口,输出可解释的事件状态;一种潜在接口,通过阶梯式解码在交错的Transformer深度间传递连续的CoT状态。一个冻结的潜在到文本重建目标为潜在表示提供了语义锚点。离线两步规划评估中,X-Planner在BERTScore-F1和基于评判者的总体得分上均位列四个评估模型中的第二。在真实机器人实验中,分别优于评估的基线。这些结果表征了规划文本质量和下游执行。

英文摘要

Task planning bridges high-level instructions and executable behavior in long-horizon manipulation, yet modern Vision-Language-Action (VLA) systems often leave this intermediate structure implicit. Existing chain-of-thought (CoT) planners also tend to rely on coarse task-level annotations or serialize long reasoning traces token by token. We present X-Planner, a planning front-end that addresses both the supervision and representation of embodied reasoning. Our planning data combine Ego, UMI, and teleoperation under a hierarchy granularity with source-dependent annotation depth. Takeover-time annotations and human-designed failures supervise ongoing error recognition. On the model side, a shared VLM backbone exposes two event-structured plan forms: a discrete interface that emits interpretable event states and a latent interface that relays continuous CoT states across staggered Transformer depths through Staircase Decoding. A frozen latent-to-text reconstruction objective provides a semantic anchor for the latent representation. Offline two-step planning evaluation places X-Planner second among four evaluated models on both BERTScore-F1 and a judge-based Overall score. In real-robot experiments, respectively, outperforming the evaluated baselines. These results characterize planning-text quality and downstream execution.

Commentshttps://github.com/X-Square-Robot/Xplanner

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑