RA-VLA:用于测试时自适应的检索增强视觉-语言-动作模型
RA-VLA: Retrieval-Augmented VLA for Test-Time Adaptation
浏览论文内容
中文总结 AI 辅助
该研究针对VLA模型面对新任务分布时的脆弱性问题,提出RA-VLA检索增强VLA框架,在LIBERO基准和UR5e环境中验证其可提升机器人操纵的任务适应成功率与计算效率。
中文摘要 AI 辅助
视觉-语言-动作(VLA)模型为通用机器人操纵提供了通用基础,但面对新任务分布时表现出显著的脆弱性。虽然上下文模仿学习(ICIL)提供了无需训练的替代方案,但现有框架存在适应瓶颈,阻碍了专家上下文向可执行动作的有效转换,这种失败源于浅层检索机制和将策略锚定在预训练先验上的固有行为惯性。为解决这些限制,我们提出RA-VLA,一种检索增强VLA框架,其整合了与行为对齐的上下文检索和基于基础的执行流水线。通过在可扩展架构中强制忠实遵循功能线索,RA-VLA在保持推理效率的同时促进无缝任务适应。我们在LIBERO基准和真实世界UR5e环境中的实证评估表明,RA-VLA实现了更优的成功率和计算效率,为无需训练的机器人适应建立了稳健框架。
英文摘要
Vision-Language-Action (VLA) models provide a versatile foundation for general robotic manipulation, yet they exhibit significant brittleness when confronted with novel task distributions. While In-Context Imitation Learning (ICIL) offers a training-free alternative, existing frameworks suffer from an adaptation bottleneck that hinders the effective translation of expert context to executable actions. This failure originates from superficial retrieval mechanisms and an inherent behavioral inertia that anchors the policy to its pre-trained priors. To address these limitations, we present RA-VLA, a retrieval-augmented VLA framework that integrates behavior-aligned context retrieval with a grounded execution pipeline. By enforcing faithful adherence to functional cues within a scalable architecture, RA-VLA facilitates seamless task adaptation while preserving inference efficiency. Our empirical evaluations across the LIBERO benchmark and a real-world UR5e environment demonstrate that RA-VLA achieves superior success rates and computational efficiency, establishing a robust framework for training-free robotic adaptation.
发表机构
- POSTECH(浦项科技大学)
- Department of Computer Science and Engineering, POSTECH(浦项科技大学计算机科学与工程学院)
- Graduate School of Artificial Intelligence, POSTECH(浦项科技大学人工智能研究生院)
机构由 AI 辅助整理,请以论文原文为准。