发表机构
Research Institute of Tsinghua University in Shenzhen; Everwise-Tech Co., Ltd.; Tongji University; Fudan University; Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳研究院; 恒智慧科技有限公司; 同济大学; 复旦大学; 清华大学深圳国际研究生院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对冻结 VLA 策略长 horizon 操作的误差问题,提出无需训练的 RTCF 框架,通过渐进式记忆对齐和低频残差转移,在 LIBERO 实验中提升了操作成功率且延迟极低。
AI 中文摘要
冻结的视觉-语言-动作(VLA)策略会生成时间上延伸的动作块,但长 horizon 操作仍易受累积执行误差和任务阶段间视觉混叠的影响。成功的 rollout 会提供有用的校正证据,然而当前帧检索可能返回进度不匹配的动作,而直接重放或时域融合会覆盖策略提议的反应结构。我们提出 Retrieve in Time, Correct in Frequency(RTCF),这是一种无需训练的测试时校正框架,能以极低的模型侧成本提升冻结 VLA 的性能。RTCF 区分要检索的经验和要转移的动作部分。渐进式记忆对齐(PMA)通过增量更新的单调边界,将不断增长的视觉执行历史与完整成功轨迹进行因果对齐,无需阶段标签即可联合识别相关记忆和当前对齐的记忆位置。从对齐的动作块中,RTCF 在运动通道上转移逐元素裁剪的低频残差,高频分量和夹爪决策则保留自冻结策略。在四个 LIBERO 套件及每个条件下 2000 个 episode 的实验中,RTCF 将总成功率从 86.4% 提升至 88.4%,并将 LIBERO-Long 的成功率从 61.6% 提升至 68.6%。这些提升无需参数更新、重复 VLA 推理或额外 GPU 资源:校正可在单次策略调用后在客户端 CPU 上执行,每个动作块的中位延迟总和仅为 10.99 毫秒。
英文摘要
Frozen vision-language-action (VLA) policies generate temporally extended action chunks, but long-horizon manipulation remains vulnerable to accumulated execution error and visual aliasing across task stages. Successful rollouts provide useful corrective evidence, yet current frame retrieval can return progress-misaligned actions,while direct replay or time-domain fusion can overwrite the reactive structure of the policy proposal. We introduce Retrieve in Time, Correct in Frequency (RTCF), a training-free test-time correction framework that improves frozen VLA performance with low model-side overhead.RTCF separates which experience to retrieve from which part of its action to transfer. Progressive Memory Alignment (PMA) causally aligns the growing visual execution history with complete successful trajectories through incrementally updated monotonic frontiers, jointly identifying a relevant memory and the current aligned memory position without stage labels. From the aligned action chunk,RTCF transfers a coefficient-wise-clipped low-frequency residual on motion channels. Higher-frequency components and gripper decisions remain inherited from the frozen policy. Across four LIBERO suites and 2,000 episodes per condition, RTCF raises aggregate success from 86.4% to 88.4% and improves LIBERO-Long from 61.6% to 68.6%.These gains require no parameter updates, repeated VLA inference, or additional GPU resources: correction can be performed on the client CPU after a single policy invocation, and the median latencies sum to only 10.99 ms per action chunk