arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

拓扑信息视觉提示用于视觉语言动作策略

Topology-Informed Visual Prompting For Vision Language Action Policies

Haoyang Wu, Abhinav Kumar, Dmitry Berenson

arXiv 2609.23944首次发表:更新:

发表机构

University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出拓扑引导的视觉提示框架,利用仿真规划和拓扑签名增强VLA策略,在复杂障碍物操控任务中优于基线,硬件成功率提升40%。

AI 中文摘要

视觉语言动作(VLA)策略在处理具有复杂障碍物几何形状的操控任务时,由于部分可观测性,可能会遇到困难。这些复杂的几何形状可能导致相似的视觉观察或机器人配置需要定性不同的动作,这种区别可以通过拓扑签名进行量化。虽然对环境几何形状和物体状态有完全了解的运动规划器可以在规划中考虑这些签名,但在部署时这些信息通常是未知的。为了解决这个问题,我们提出了一种拓扑引导的视觉提示框架,该框架利用基于仿真的规划来增强名义演示数据集,并在部署时提供基于视觉的指导。我们的方法使用高斯链接积分拓扑签名表示来捕捉环境的重要拓扑属性。利用环境仿真近似中的特权几何信息,我们用将系统移动到已演示签名的轨迹来增强VLA微调数据集,并从新配置恢复任务执行。在同一数据集上微调视觉语言模型(VLM),以从实时相机观察中预测签名和末端执行器航点,这些航点被渲染为观察上的视觉提示以引导VLA。在三个模拟双臂任务和一个真实世界的箱子拾取任务中,我们的方法优于仅在名义演示上微调的VLA和能够从观察中移除拓扑相关信息的VLM提示基线。在硬件上,它在任务成功率上超过最强基线40%。项目网站:此https URL。

英文摘要

Vision-language-action (VLA) policies can struggle with manipulation tasks with complex obstacle geometries due to partial observability. These complex geometries can lead to similar visual observations or robot configurations requiring qualitatively different actions, a distinction that can be quantified using topological signatures. While motion planners with full knowledge of environment geometries and object states can reason about these signatures in planning, this information is often not known at deployment. To address this issue, we present a topology-guided visual-prompting framework that uses simulation-based planning to augment a nominal demonstration dataset and provides vision-based guidance at deployment. Our method uses a Gauss-Linking-Integral topological signature representation to capture important topological properties of the environment. Using privileged geometry information from a simulation approximation of our environment, we augment a VLA fine-tuning dataset with trajectories that move the system to a demonstrated signature and, from the new configuration, resume task execution. A vision-language model (VLM) is fine-tuned on the same dataset to both predict signatures from live camera observations and predict end-effector waypoints, which are rendered as visual prompts on the observations to guide the VLA. Across three simulated bimanual tasks and a real-world box pickup task, our method outperforms a VLA fine-tuned only on nominal demonstrations and a VLM-prompting baseline that can remove topology-relevant information from observations. On hardware, it exceeds the strongest baseline by 40% in task success. Project website: https://topology-vla.github.io.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑