发表机构
Suzhou Institute for Advanced Research, University of Science and Technology of China; School of Biomedical Engineering, Division of Life Sciences and Medicine, University of Science and Technology of China(中国科学技术大学苏州高等研究院; 中国科学技术大学生命科学与医学部生物医学工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GRAFT是一种高效在线VLA适应框架,通过特定视角视觉锚点聚焦任务相关线索,在四项生物医学操作任务中提升25个百分点成功率并降低计算开销。
AI 中文摘要
预训练的视觉-语言-动作(VLA)策略为机器人操作提供了强大的先验,但将其在线适应到精细生物医学任务仍具挑战性。任务成功往往依赖于细微的、视角相关的视觉线索,而任务级奖励几乎无法指导哪些区域重要,因此难以从有限的真实机器人交互中学习任务相关的视觉 grounding。在线适应还受到VLA推理和基于回放的更新的计算成本限制。我们提出GRAFT(面向快速任务学习的 grounded 强化适应),这是一种通过 grounded 感知实现高效在线VLA适应的框架。GRAFT利用区域级监督学习特定视角的视觉锚点,将感知聚焦于任务相关的局部线索,无需在部署时生成区域提议。它还将单步动作生成与缓存的视觉-语言前缀重用相结合,以加速在线学习。在四项生物医学操作任务中,GRAFT在匹配的适应预算下将成功率提高了25个百分点,同时降低了在线策略更新的计算开销。
英文摘要
Pretrained vision-language-action (VLA) policies provide strong priors for robot manipulation, yet adapting them online to fine-grained biomedical tasks remains challenging. Task success often hinges on subtle, view-dependent visual cues, while task-level rewards provide little guidance about which regions matter, making it difficult to learn task-relevant visual grounding from limited real-robot interaction. Online adaptation is further constrained by the computational cost of VLA inference and replay-based updates. We introduce GRAFT (Grounded Reinforcement Adaptation for Fast Task Learning), a framework for efficient online VLA adaptation through grounded perception. GRAFT uses region-level supervision to learn view-specific visual anchors that focus perception on task-relevant local cues without requiring region proposals at deployment. It further combines single-step action generation with cached visual-language prefix reuse to accelerate online learning. Across four biomedical manipulation tasks, GRAFT improves success rates by 32.5 percentage points under matched adaptation budgets, while reducing the computational overhead of online policy updates.