发表机构
Wuhan University; Institute of Technological Sciences, Wuhan University(武汉大学; 武汉大学技术科学研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ICI-VLA提出一种训练与检索框架,通过时空对齐示范实现VLA模型少样本测试时自适应,在LIBERO和RoboTwin 2.0上分别取得97.7%和60.4%的平均成功率,并超越基线。
AI 中文摘要
视觉-语言-动作(VLA)策略通常通过额外的梯度更新来适应新的操作环境,这在任务特定数据或计算资源稀缺时限制了快速部署。我们提出了ICI-VLA,一个训练和检索框架,通过上下文示范赋予文本-动作视觉语言模型(VLM)少样本测试时自适应能力。与基于动作特定多模态融合的主流VLA设计不同,ICI-VLA保留了原生的文本生成接口。ICI-VLA仅在离线训练期间更新其参数;在推理时,策略保持固定,并根据检索到的微示范来条件化动作生成。该框架将长轨迹分解为短的、带语义标签的示例,并训练一个RD-Encoder,使用动态时间规整(DTW)挖掘的正样本,使检索到的上下文与当前子任务的阶段和几何形状对齐。我们进一步引入了目标动作掩蔽(Target Action Masking),这是一种上下文损坏目标,旨在减少直接动作复制并增加对当前观测的依赖。ICI-VLA在LIBERO上达到平均成功率97.7%,在RoboTwin 2.0上达到60.4%,超过RoboTwin 2.0上报告的最高基线平均值19.3个百分点。它还在四项物理任务中达到83.2%的成功率。这些结果表明,固定的VLA策略可以通过在测试时条件化于时空对齐的示范而受益。
英文摘要
Vision-Language-Action (VLA) policies are commonly adapted to new manipulation settings through additional gradient updates, which limits rapid deployment when task-specific data or compute is scarce. We present ICI-VLA, a training and retrieval framework that equips a text-action VLM with few-shot test-time adaptation through in-context demonstrations. Unlike mainstream VLA designs based on action-specific multimodal fusion, ICI-VLA retains the native text-generation interface. ICI-VLA updates its parameters only during offline training; at inference, the policy remains fixed and conditions action generation on retrieved micro-demonstrations. The framework decomposes long trajectories into short, semantically labeled examples and trains an RD-Encoder with positives mined by Dynamic Time Warping (DTW), aligning the retrieved context with the phase and geometry of the current subtask. We further introduce Target Action Masking, a context-corruption objective designed to reduce direct action copying and increase reliance on the current observation. ICI-VLA reaches average success rates of 97.7% on LIBERO and 60.4% on RoboTwin 2.0, exceeding the highest reported baseline average on RoboTwin 2.0 by 19.3 percentage points. It also achieves 83.2% across four physical tasks. These results indicate that a fixed VLA policy can benefit from conditioning on spatiotemporally aligned demonstrations at test time.
Comments9 pages, 6 figures