发表机构
KU Leuven(鲁汶大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出空间语言建模方法,用自回归Transformer统一策略学习与状态预测,在Push-T仿真和真实机器人任务上,其真实机器人性能优于基线,还可预测场景状态。
AI 中文摘要
学习动作如何改变场景几何结构可为目标导向的操纵提供互补监督。我们提出空间语言建模(Spatial Language Modeling),其用离散坐标和语义标记的共享词汇表表示场景轮廓、目标、动作目标及未来状态。特定任务的语法将这些元素组织成空间序列,使单个自回归Transformer通过共同的下一个标记目标学习动作生成和动作条件下的状态预测。我们从头开始训练该模型,先使用随机交互转换预训练,再在专家演示上进行动作与状态的联合训练。预训练期间,记录的动作坐标为后续状态预测提供条件,且不被纳入预测损失。控制阶段,模型仅解码可执行的动作目标,并以新观测到的状态更新自身历史。我们在仿真环境的Push-T任务及真实机器人上评估该方法,该模型在仿真中取得有竞争力的性能,且在真实机器人上的任务成功率和目标覆盖率均高于所评估的策略基线。训练消融实验显示,动作与状态序列的联合训练可提升控制效果,随机交互预训练还能进一步带来性能增益。给定提供的动作轨迹,同一模型还能预测连续的场景状态,捕捉推动动作的几何效应。
英文摘要
Learning how actions change scene geometry can provide complementary supervision for goal-directed manipulation. We introduce Spatial Language Modeling, which represents scene contours, goals, action targets, and future states with a shared vocabulary of discrete coordinates and semantic tokens. A task-specific grammar organizes these elements into spatial sequences, allowing one autoregressive Transformer to learn action generation and action-conditioned state prediction through a common next-token objective. We train the model from scratch using random-play transition pretraining followed by joint action and state training on expert demonstrations. During pretraining, recorded action coordinates condition subsequent state predictions and are excluded from the prediction loss. During control, the model decodes only executable action targets and updates its history with newly observed states. We evaluate the approach on Push-T in simulation and on a real robot. The model achieves competitive simulation performance and higher task success and target coverage than the evaluated real-robot policy baselines. Training ablations show improved control with joint action and state sequences, with further gains from random-play pretraining. Given supplied action trajectories, the same model also predicts successive scene states, capturing the geometric effects of pushing.