动作草图生成器:通过视觉草图实现从推理到动作的长周期机器人操作
Action-Sketcher: From Reasoning to Action via Visual Sketches for Long-Horizon Robotic Manipulation
- State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(1 多媒体信息处理国家重点实验室,计算机科学学院,北京大学)
- Beijing Academy of Artificial Intelligence(2 北京人工智能研究院)
- University of Sydney(3 新南威尔士大学)
- Institute of Automation, Chinese Academy of Sciences(4 中国科学院自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
动作草图生成器通过视觉草图实现从推理到动作的长周期机器人操作,结合视觉-语言-动作框架提升任务执行的鲁棒性和可解释性。
AI中文摘要:
长周期机器人操作在现实世界部署中越来越重要,要求在复杂布局中进行空间歧义消除,并在动态交互下保持时间韧性。然而,现有的端到端和分层视觉-语言-动作(VLA)策略通常依赖纯文本提示,同时保持计划意图隐含,这会削弱在杂乱或信息不足的场景中的指称联系,阻碍通过闭环交互有效分解长周期目标,并限制因果解释,因为行动选择背后的理由被遮蔽。为了解决这些问题,我们首先引入了视觉草图,一种不可信的视觉中间表示,它在机器人的当前视图中渲染点、框、箭头和带类型的关系,以外部化空间意图,连接语言到场景几何。在视觉草图的基础上,我们提出了动作草图生成器,一种VLA框架,它在一个循环的看-想-草图-行动工作流程中运行,由适应性令牌门控策略协调推理触发、草图修订和行动发布,从而支持反应性修正和人机交互,同时保持实时行动预测。为了实现可扩展的训练和评估,我们整理了多样化的语料库,包含交错的图像、文本、视觉草图监督和行动序列,并通过结合交错序列对齐用于模态统一、语言到草图一致性用于精确语言联系、以及带有草图到行动强化学习的模仿学习的多阶段课程食谱来训练动作草图生成器。在杂乱场景和多对象任务上的大量实验,包括仿真和现实任务,显示了改进的长周期成功率、更强的动态场景变化鲁棒性和通过可编辑草图和分步计划增强的可解释性。项目网站:https://action-sketcher.github.io
英文摘要:
Long-horizon robotic manipulation is increasingly important for real-world deployment, requiring spatial disambiguation in complex layouts and temporal resilience under dynamic interaction. However, existing end-to-end and hierarchical Vision-Language-Action (VLA) policies often rely on text-only cues while keeping plan intent latent, which undermines referential grounding in cluttered or underspecified scenes, impedes effective task decomposition of long-horizon goals with close-loop interaction, and limits causal explanation by obscuring the rationale behind action choices. To address these issues, we first introduce Visual Sketch, an implausible visual intermediate that renders points, boxes, arrows, and typed relations in the robot's current views to externalize spatial intent, connect language to scene geometry. Building on Visual Sketch, we present Action-Sketcher, a VLA framework that operates in a cyclic See-Think-Sketch-Act workflow coordinated by adaptive token-gated strategy for reasoning triggers, sketch revision, and action issuance, thereby supporting reactive corrections and human interaction while preserving real-time action prediction. To enable scalable training and evaluation, we curate diverse corpus with interleaved images, text, Visual Sketch supervision, and action sequences, and train Action-Sketcher with a multi-stage curriculum recipe that combines interleaved sequence alignment for modality unification, language-to-sketch consistency for precise linguistic grounding, and imitation learning augmented with sketch-to-action reinforcement for robustness. Extensive experiments on cluttered scenes and multi-object tasks, in simulation and on real-world tasks, show improved long-horizon success, stronger robustness to dynamic scene changes, and enhanced interpretability via editable sketches and step-wise plans. Project website: https://action-sketcher.github.io