arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23275cs.RO

SCULPT-VLA:通过分阶段动作接地学习结构化控制

SCULPT-VLA: Learning Structured Control through Staged Action Grounding

  • South China University of Technology(华南理工大学)
  • Yuanwu Technology(元武科技)

机构由 AI 辅助整理,请以论文原文为准。

Wenbo Li, Yiteng Chen, Wei Zhang, Wenhao Li, Jun Yang, Qingyao Wu

AI总结:

SCULPT-VLA通过分阶段动作接地,将结构化状态学习与连续控制细化分离,在多个基准和物理机器人任务上超越基线,实现更高成功率。

AI中文摘要:

视觉-语言-动作(VLA)策略越来越多地在动作标签之外融入结构化中间监督。然而,指定中间表示应编码什么内容,却留下了动作预测如何学习依赖该表示的问题。我们提出SCULPT-VLA,一种通过分阶段动作接地学习结构化控制的策略。其动作条件状态由任务进展、场景动态和空间接地的互补因素组成。训练首先利用教师脚手架形成这些因素,然后在脚手架输入撤除时,通过它们的组合来接地粗粒度动作预测。随后恢复直接感知访问以进行连续细化,将学习到的状态与感知细节相结合。该课程将学习在结构上条件化动作与细化连续控制分离开来。部署既不需要教师,也不需要离散动作自回归。SCULPT-VLA在LIBERO、SimplerEnv-WidowX和RoboTwin 2.0 Full上比共享骨干基线实现了更高的平均成功率。在SimplerEnv-WidowX上,最终成功率为83.5%,而第二阶段动作学习直接访问视觉和语言时为71.3%。在四个物理机器人任务中,测试分布偏移下的平均成功率达到58.1%,而π0.5为45.6%。训练消融和因素级干预支持了分阶段设计,并表明在恢复直接感知访问后,学习到的状态继续对控制做出贡献。

英文摘要:

Vision-language-action (VLA) policies increasingly incorporate structured intermediate supervision beyond action labels. Yet specifying what an intermediate representation should encode leaves open how action prediction learns to depend on it. We introduce \textbf{SCULPT-VLA}, a policy that learns structured control through staged action grounding. Its action-conditioning state comprises complementary factors for task progression, scene dynamics, and spatial grounding. Training first forms these factors with teacher scaffolds, then grounds coarse action prediction through their composition as scaffold inputs are withdrawn. Direct perceptual access is subsequently restored for continuous refinement, combining the learned state with perceptual detail. The curriculum separates learning to condition actions on structure from refining continuous control. Deployment requires neither teachers nor discrete-action autoregression. SCULPT-VLA achieves higher average success than shared-backbone baselines on LIBERO, SimplerEnv-WidowX, and RoboTwin 2.0 Full. On SimplerEnv-WidowX, final success is 83.5\%, versus 71.3\% when Stage-II action learning directly accesses vision and language. Across four physical robot tasks, average success under the tested distribution shifts reaches 58.1\%, compared with 45.6\% for $π_{0.5}$. Training ablations and factor-wise interventions support the staged design and show that the learned state continues to contribute to control after direct perceptual access is restored.

补充信息

↑