发表机构
X Square Robot Team(X Square Robot 团队)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
X-Planner提出事件结构化任务规划前端,结合离散与潜在接口及阶梯式解码,改善具身智能的推理监督与表示,在离线评估中排名第二并优于真实机器人基线。
AI 中文摘要
任务规划在长时程操作中连接高层指令与可执行行为,然而现代的视觉-语言-动作(VLA)系统往往使这一中间结构隐式化。现有的思维链(CoT)规划器也倾向于依赖粗略的任务级注释,或逐词元地序列化冗长的推理轨迹。我们提出X-Planner,一个规划前端,旨在同时解决具身推理的监督与表示问题。我们的规划数据在层级粒度下结合了Ego、UMI和遥操作,并具有依赖数据源的注释深度。接管时注释和人为设计的失败用于监督持续的错误识别。在模型方面,共享的VLM主干暴露了两种事件结构化规划形式:一种离散接口,输出可解释的事件状态;一种潜在接口,通过阶梯式解码在交错的Transformer深度间传递连续的CoT状态。一个冻结的潜在到文本重建目标为潜在表示提供了语义锚点。离线两步规划评估中,X-Planner在BERTScore-F1和基于评判者的总体得分上均位列四个评估模型中的第二。在真实机器人实验中,分别优于评估的基线。这些结果表征了规划文本质量和下游执行。
英文摘要
Task planning bridges high-level instructions and executable behavior in long-horizon manipulation, yet modern Vision-Language-Action (VLA) systems often leave this intermediate structure implicit. Existing chain-of-thought (CoT) planners also tend to rely on coarse task-level annotations or serialize long reasoning traces token by token. We present X-Planner, a planning front-end that addresses both the supervision and representation of embodied reasoning. Our planning data combine Ego, UMI, and teleoperation under a hierarchy granularity with source-dependent annotation depth. Takeover-time annotations and human-designed failures supervise ongoing error recognition. On the model side, a shared VLM backbone exposes two event-structured plan forms: a discrete interface that emits interpretable event states and a latent interface that relays continuous CoT states across staggered Transformer depths through Staircase Decoding. A frozen latent-to-text reconstruction objective provides a semantic anchor for the latent representation. Offline two-step planning evaluation places X-Planner second among four evaluated models on both BERTScore-F1 and a judge-based Overall score. In real-robot experiments, respectively, outperforming the evaluated baselines. These results characterize planning-text quality and downstream execution.
Commentshttps://github.com/X-Square-Robot/Xplanner