SpaceVLA:用于机器人操作的空间 grounding VLA,支持用户编写的抓取与放置锚点
SpaceVLA: Spatially Grounded VLA for Robotic Manipulation with User-Authored Grasp and Place Anchors
浏览论文内容
中文总结 AI 辅助
本研究针对VLA模型操作时缺乏显式空间意图的问题,提出视觉意图锚点的XR流程,通过微调OpenVLA-7B的SpaceVLA模型,在Unity试验中实现91.25%抓取成功率,完成机器人抓取放置任务。
中文摘要 AI 辅助
视觉-语言-动作(VLA)模型可遵循语言指令,但往往缺乏操作所需的显式空间意图。我们提出视觉意图锚点,这是一种XR流程,允许用户指定抓取和放置区域,并将其渲染为图像空间叠加层,供VLA控制使用。我们收集200个Unity抓取-放置演示,并使用LoRA在时间上 subsampled 的带注释观测上微调OpenVLA-7B。该策略从标记的RGB观测和语言中预测tokenized的7自由度增量动作。我们在闭环Unity试验中评估该策略,实现91.25%的抓取成功率,平均抓取误差为0.5厘米,平均放置误差为0.7厘米。
英文摘要
Vision-language-action (VLA) models follow language commands but often lack explicit spatial intent for manipulation. We present Visual Intent Anchors, an XR pipeline that lets users specify grasp and placement regions and renders them as image-space overlays for VLA control. We collect 200 Unity pick-and-place demonstrations and fine-tune OpenVLA-7B with LoRA on temporally subsampled annotated observations. The policy predicts tokenized 7-DoF incremental actions from marked RGB observations and language. We evaluate the policy in closed-loop Unity trials, achieving a grasp success rate of 91.25% and mean grasp and placement errors of 0.5 cm and 0.7 cm, respectively.