检索前先复用:诊断具身多模态策略测试时增强的余量与互补性
Reuse Before You Retrieve: Diagnosing Headroom and Complementarity for Test-Time Augmentation of Embodied Multimodal Policies
浏览论文内容
中文总结 AI 辅助
该研究提出通过可恢复余量与检索互补性两个因素,为具身多模态冻结 VLA 策略的测试时增强选择采样或检索干预,在 LIBERO 上成功提升最高 21.0 个百分点,且可迁移至其他机器人与环境。
中文摘要 AI 辅助
冻结的视觉-语言-动作(VLA)策略正日益通过采样额外策略行为或引入外部演示在测试阶段得到改进。然而,对于部署的策略究竟需要哪种干预措施,目前几乎没有指导原则。额外采样仅在策略的随机 rollout 中已存在更好行为且可被识别时才有用,而当相关动作先验未被策略可靠表征时,检索最为有用。我们通过两个可测量因素——可恢复余量(recoverable headroom)和检索互补性(retrieval complementarity)——来研究这一决策,二者分别表征已有多少可用行为可被恢复,以及外部动作先验是否填补了可测量的差距。我们在可重试或并行执行的场景下评估了一个 episode 级重试选择器,同时评估了跨多个冻结 VLA 策略和环境的检索效果。该选择器在 LIBERO 数据集上针对所有测试的 VLA 骨干网络,始终恢复了大量潜在能力,成功提升最高达 21.0 个百分点,且与可恢复余量紧密对应。它还能迁移至不同机器人和模拟器,在观测退化时仍保持有效;对自回归 OpenVLA 的实验则阐明了可用余量与对候选 rollout 排序能力之间的区别。检索表现不同,会在测量到最大动作先验差距的策略上进行改进,且与选择结合时能提供进一步增益。这些结果共同为测试时增强机会的表征提供了经验基础,方法是将可从冻结策略中恢复的能力与可能需要从外部引入的行为先验区分开来。
英文摘要
Frozen vision-language-action (VLA) policies are increasingly improved at test time by sampling additional policy behaviors or introducing external demonstrations. Yet there is little guidance for deciding which intervention a deployed policy actually needs. Additional sampling is useful only when better behavior already exists within the policy's stochastic rollouts and can be identified, whereas retrieval is most useful when the relevant action prior is not reliably represented by the policy. We study this decision through two measurable factors, recoverable headroom and retrieval complementarity, which characterize how much useful behavior is already available to recover and whether an external action prior fills a measurable gap. We evaluate an episode-level retry selector under retryable or parallel execution, together with retrieval across multiple frozen VLA policies and environments. The selector consistently recovers substantial latent capability across all tested VLA backbones on LIBERO, with gains of up to 21.0 success-rate points that closely track recoverable headroom. It also transfers to a different robot and simulator and remains effective under degraded observations, while experiments with autoregressive OpenVLA illustrate the distinction between available headroom and the ability to rank candidate rollouts. Retrieval behaves differently, improving the policy with the largest measured action-prior gap and providing further gains when combined with selection. Together, these results provide an empirical basis for characterizing test-time augmentation opportunities by separating capability that can be recovered from the frozen policy from behavioral priors that may need to be introduced externally.
发表机构
- KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。