arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

将意图与轨迹解耦:面向世界动作模型的表示演绎框架

Decoupling Intention from Trajectory: A Representational Deduction Framework for World Action Models

Xiangkai Ma, Yue Ma, Junjie Wang, Sheng Xu, Mingyang Li, Han Zhang, Yuzheng Zhuang, Wenzhong Li, Zhihao Yuan

arXiv 2608.06994首次发表:更新:

发表机构

NJU; HKUST; CUHK-SZ; THU; Joy Future Academy, JD(南京大学; 香港科技大学; 香港中文大学(深圳); 清华大学; 京东探索研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有世界动作模型的表示纠缠问题,提出PILOT的表示演绎框架,解耦运动语义与轨迹细节,提升复杂机器人操纵任务的性能与可解释性,兼具少样本微调及架构迁移的可扩展性。

AI 中文摘要

世界动作模型(World Action Models, WAMs)旨在构建统一架构,用于理解世界状态演化并指导生成式运动规划。然而,现有视觉分支仅聚焦于预测静态视觉观测,未捕捉运动交互下的世界状态演化潜在转移信息,导致动作模型内高级物理状态演化与低级动作轨迹生成的表示纠缠,形成结构瓶颈,削弱了世界演化建模对动作生成的预测能力。本文提出PILOT(Physical Inference for Latent Optimized Trajectories,潜在优化轨迹的物理推理),其核心表示演绎(Representational Deduction, RD)通过将运动思维链(Chain-of-Thought, CoT)引导作为模型原生能力,弥合上述差距。具体而言,RD旨在促使动作分支显式建模潜在状态转移标记,将其作为推理空间中的CoT保留,以指导细粒度运动轨迹。实验表明,RD不仅显著提升了WAM在复杂机器人操纵任务中的成功率与泛化能力,还通过将高级运动语义与低级轨迹细节解耦增强了模型的物理解释性;此外,RD引入的丰富状态转移监督信号有效缓解了动作生成中的稀疏监督问题,可作为高效的少样本真实机器人微调策略,且在向主流WAM架构迁移时展现出优异的可扩展性。

英文摘要

World Action Models (WAMs) aim to construct a unified architecture capable of understanding world state evolution and guiding to generative motion planning. However, existing visual branches focus on predicting static visual observation, rather than reflecting potential transition information that captures the evolution of world states under motion interactions. This leads to representational entanglement between high-level physical condition evolution and low-level action trajectory generation within the Action Model, creating a structural bottleneck while weakening the predictive capability of world evolution modeling for action generation. We propose PILOT (Physical Inference for Latent Optimized Trajectories), whose core Representational Deduction (RD) bridges this gap by integrating motion thought-of-chain (CoT) guidance as a native model capability. Specifically, RD aims to encourage the action branch to explicitly model potential state transition tokens, which are retained as CoT in the reasoning space to guide fine-grained motion trajectory. Experiments demonstrate that RD not only significantly improves the success rate and generalization ability of WAMs in complex robotic manipulation tasks but also enhances the model's physical interpretability by decoupling high-level motion semantics from low-level trajectory details. Furthermore, the abundant state transition supervision signals introduced by RD effectively alleviate the sparse supervision in action generation, enabling it to serve as an efficient few-shot real-robot fine-tuning strategy and demonstrating superior scalability for migration to mainstream WAM architectures.

CommentsAt the request of our institution, we are withdrawing this preprint pending completion of the institutional clearance process for public release

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑