RedFlow:将失败重定向为流匹配VLA策略的动作级修正
RedFlow: Redirect Failure into Action-level Corrections for Flow-matching VLA Policy
查看机构详情
- The Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
RedFlow是将失败转化为动作级修正监督的离线RL框架,在LIBERO基准及真实操纵任务上,其真实成功率达74.7%,优于现有离线RL基线,样本量仅为强在线方法的十分之一。
中文摘要 AI 辅助
流匹配视觉-语言动作(VLA)策略在机器人操纵领域展现出强大潜力,但部署过程中常因分布偏移引发误差累积。离线强化学习(RL)为利用rollout数据改进部署策略提供了可行方案,不过现有方法要么忽略失败数据,要么仅在轨迹层面利用这类数据,导致学习效率低下且误差持续存在。我们提出RedFlow,这是一种细粒度离线RL框架,可将失败经验重定向为流匹配VLA策略的动作级修正监督信号。RedFlow包含两个核心组件:一是**上下文感知修正匹配机制**,用于识别引发失败的动作,并从相似上下文的成功案例中检索替代方案作为修正目标;二是**自适应重定向目标函数**,该函数会同时强化成功动作、抑制不良动作,并将可恢复的失败案例重定向至修正目标。通过将成功与失败经验均转化为密集监督信号,RedFlow可从混合质量的数据中实现鲁棒的恢复学习。在LIBERO基准测试及三个真实世界操纵任务上开展的实验表明,RedFlow的性能始终优于现有最优离线RL基线,将真实世界任务的成功率从56.7%提升至74.7%;同时,它还能与PPO、GRPO、DDPO等强在线方法表现相当,且所需训练样本量约为这些方法的十分之一。
英文摘要
Reinforcement learning (RL) can improve Vision-Language-Action (VLA) policies from deployment experience, but reward- and preference-based RL primarily identifies desirable behaviors without specifying how to correct failed actions, underutilizing failure trajectories and limiting sample efficiency. Can such corrections be derived from fixed rollouts? Our key insight is that rollouts with different outcomes may contain action chunks executed in similar states, enabling higher-quality chunks to provide locally supported corrective references. Building on this insight, we introduce \textbf{RedFlow}, an offline post-training method for flow-matching VLA policies. \emph{Execution-Context Matching} groups chunks using a compact representation of estimated task progress and robot proprioception. \emph{Quality-Guided Action Redirection} assigns signed chunk-quality scores and aggregates higher-quality chunks into corrective targets, reinforcing high-quality chunks, suppressing low-quality chunks, and redirecting correctable chunks toward their targets. RedFlow requires neither external HIL corrections nor online data collection during post-training. Across four LIBERO suites, RedFlow improves average success from 56.2\% to 68.2\%, outperforming the strongest evaluated offline baseline, AWR (62.3\%), by 5.9 points. Across three real-robot tasks, it improves average success from 56.7\% to 74.7\%. On LIBERO-Spatial, RedFlow reaches 75.8\% success with 1{,}536 fixed rollouts, while the evaluated online methods require 8.7--16$\times$ as many fresh post-training rollouts to reach the same threshold.