arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

XGenAct:通过跨任务生成实现几何增强的世界动作模型

XGenAct: Geometry-Enhanced World Action Models through Cross-Task Generation

Tingting Du, Ziyao Wang, Guoheng Sun, Ang Li

arXiv 2610.03516首次发表:更新:

发表机构

University of Wisconsin–Madison; University of Maryland, College Park(威斯康星大学麦迪逊分校; 马里兰大学学院公园分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

XGenAct提出将深度、法向量、分割等空间信息编码为RGB视频,用单一视频扩散变换器统一预测,在RLBench上闭环成功率显著优于基线。

AI 中文摘要

世界动作模型(WAMs)通过预测观测和动作随时间的演变,推动了机器人控制的发展。尽管取得了这些进展,基于RGB和动作的未来预测并未明确解决机器人操作所需的空间理解问题。现有的工作往往通过专门的头部或分支添加有限的一组空间预测任务,导致空间监督的范围和模型架构都显得碎片化。我们提出了XGenAct,一种世界动作模型,它将RGB观测、机器人动作、度量深度、表面法向量和功能角色分割通过确定性编解码器表示为RGB视频。通过在训练过程中采样感知和动作任务,XGenAct使用一个视频扩散变换器和一个目标函数来学习跨这些空间的时间预测,而无需特定模态的学习头部。在保留的RLBench任务上,结构化感知训练相比仅使用RGB训练,提高了平均闭环成功率,并且XGenAct在五项任务的外部比较中达到了52%的成功率,而评估的最强基线仅为26%。此外,与先生成RGB再应用冻结感知专家的评估流程相比,XGenAct能更准确地预测未来的深度和分割结果。

英文摘要

World action models (WAMs) have advanced robot control by predicting how observations and actions evolve over time. Despite this progress, RGB and action based future prediction does not explicitly address the spatial understanding needed for robot manipulation. Existing efforts often add a limited set of spatial prediction tasks through specialized heads or branches, leaving both the range of spatial supervision and the model architecture fragmented. We introduce XGenAct, a world action model that represents RGB observations, robot actions, metric depth, surface normals, and functional role segmentation as RGB videos through deterministic codecs. By sampling perception and action tasks during training, XGenAct uses one video diffusion transformer and one objective to learn temporal prediction across these spaces without modality specific learned heads. On held out RLBench tasks, structured perception training improves average closed loop success over RGB only training, and XGenAct achieves 52% success in the five task external comparison, versus 26% for the strongest evaluated baselines. It also predicts future depth and segmentation more accurately than the evaluated pipelines that generate RGB first and then apply a frozen perception expert.

Comments27 pages, including appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑