arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.08448cs.RO

利用VLA:通过记忆引导智能体将冻结的视觉语言动作模型转化为可靠的操作原语

Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents

Yixian Zhang, Huanming Zhang, Feng Gao, Xiao Li, Zhihao Liu, Yi Nie, Chunyang Zhu, Jiaxing Qiu, Yuchen Yan, Jiyuan Liu, Wenhao Tang, Jiaji Rao, Zhengru Fang, Ch… 展开作者

Yixian Zhang, Huanming Zhang, Feng Gao, Xiao Li, Zhihao Liu, Yi Nie, Chunyang Zhu, Jiaxing Qiu, Yuchen Yan, Jiyuan Liu, Wenhao Tang, Jiaji Rao, Zhengru Fang, Changxu Wei, Yu Wang, Wenbo Ding, Chao Yu

首次发表
浏览论文内容

中文总结 AI 辅助

研究语言条件下操作问题,提出Harness VLA框架,将冻结VLA与分析原语库结合,通过学习操作范围扩展预训练VLA,在多种扰动操作任务中相比基线有大幅提升。

中文摘要 AI 辅助

语言条件下的操作既需要精确的富含接触的控制,也需要对语言、场景和长视野进行稳健推理。端到端视觉语言动作(VLA)模型具备强大的局部视觉运动技能,但在部署扰动下常失败。语言模型编码智能体提供互补推理,但纯分析原语在不规则抓取等方面存在困难。我们提出Harness VLA,一个记忆增强的智能体框架,将冻结的VLA作为可重试的富含接触原语,并与少量固定分析原语库组合。该框架从任务特定执行轨迹等学习固定原语的操作范围,在不微调的情况下扩展预训练VLA。在多种扰动操作任务中,Harness VLA相比最强基线有显著提升。

英文摘要

Language-conditioned manipulation requires both precise contact-rich control and robust reasoning over language, scenes, and long horizons. End-to-end Vision-Language-Action (VLA) models provide strong local visuomotor skills, but they are trained on in-distribution task trajectories and often fail under deployment perturbations such as semantic retargeting, goal re-binding, spatial-layout shifts, and unstable local contacts. LLM coding agents provide complementary semantic and compositional reasoning, but purely analytic primitives struggle with irregular grasping, constrained placement, and articulated-object interaction. We present Harness VLA, a memory-augmented agentic framework that exposes a frozen VLA as a retryable contact-rich primitive and composes it with a small fixed library of analytic primitives for grounding, staging, transport, navigation, and release. Rather than expanding the skill library, the harness learns the operating range of these fixed primitives from task-specific execution traces, global success rules, and failure models. By lifting semantic re-grounding, non-contact execution, and VLA re-staging to the planner while reserving the frozen VLA for local contact-rich phases, Harness VLA extends pretrained VLAs beyond their original trajectory distribution without finetuning. Across perturbed tabletop, household kitchen, and clean-to-randomized bimanual manipulation, Harness VLA improves over the strongest relevant baselines by 38.6 and 25.4 percentage points on LIBERO-Pro and RoboCasa365, respectively, and reaches 58.4% on RoboTwin C2R. Real-world demonstrations on dual-Franka robots further show target redirection, grasp recovery, and new task compositions with the same frozen VLA. Code is available at https://github.com/RLinf/RPent.

发表机构

  • Tsinghua University(清华大学)
  • Striding AI(驰行人工智能公司)
  • Purdue University(普渡大学)
  • Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
  • Infinigence AI(英飞智研人工智能公司)
  • Zhongguancun Academy(中关村学院)
  • Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

↑