检索何时有帮助?视觉-语言-动作模型中上下文内适应的研究
When Does Retrieval Help? A Study of In-Context Adaptation in Vision-Language-Action Models
浏览论文内容
中文总结 AI 辅助
本研究系统比较了四种检索方法在VLA模型上下文内适应中的效果,发现无一致最优方法,随机检索也有非平凡成功率,且标准检索质量指标不能可靠预测下游性能,但相关任务演示可提供可迁移信息。
中文摘要 AI 辅助
视觉-语言-动作(VLA)模型作为通用机器人策略展现出巨大潜力,但将其适应于未见过的任务通常需要昂贵的参数更新。近期如RICL等工作通过基于当前VLA观测检索专家演示并在测试时提供额外上下文,引入了上下文内适应性。因此,这种适应的有效性关键取决于检索机制。在本工作中,我们系统研究了不同检索方法如何在RICL框架内影响检索质量和任务性能。具体而言,我们比较了四种方法:基于图像的检索、结合VLA状态的检索、使用VLA骨干网络特征的检索以及随机检索。我们的实验得出三个主要发现。第一,没有任何检索方法在任务成功率上持续优于其他方法,而令人惊讶的是,随机检索也取得了非平凡的成功率。第二,标准的检索质量诊断并不能可靠地反映下游VLA性能。第三,来自不同但相关任务的演示可以提供有用的可迁移信息。综合来看,这些结果为理解检索机制如何塑造VLA模型的上下文内学习能力及其下游任务性能迈出了初步一步,同时强调了为可靠的测试时适应需要更仔细地设计和评估检索机制。
英文摘要
Vision-language-action (VLA) models have shown strong potential as generalist robot policies, but adapting them to unseen tasks often requires costly parameter updates. Recent work such as RICL introduces in-context adaptability by retrieving expert demonstrations based on the current VLA observation and providing them as additional context at test time. The effectiveness of this adaptation therefore depends critically on the retrieval mechanism. In this work, we systematically study how different retrieval methods affect both retrieval quality and task performance within the RICL framework. Specifically, we compare four different methods: image-based retrieval, retrieval augmented with VLA's state, retrieval using features from the VLA backbone, and random retrieval. Our experiments yield three main findings. First, no retrieval method consistently dominates the others in task success, while surprisingly, random retrieval achieves a non-trivial success rate. Second, standard retrieval-quality diagnostics do not reliably reflect downstream VLA performance. Third, demonstrations from different but related tasks can provide useful transferable information. Together, these results provide an initial step toward understanding how retrieval mechanisms shape the in-context learning capability of VLA models and their downstream task performance, while highlighting the need for more careful design and evaluation of retrieval mechanisms for reliable test-time adaptation.
发表机构
- Tulane University(杜兰大学)
- Aalto University(阿尔托大学)
机构由 AI 辅助整理,请以论文原文为准。