AI 中文总结
PF-RL通过学习目标条件进度场几何,将中间进度转换为密集信用,用于VLA策略的离线改进和在线强化微调,在多个基准上超越现有方法。
AI 中文摘要
强化微调(RFT)已成为改进视觉-语言-动作(VLA)策略的一种有前景的范式,然而稀疏的任务级结果仅能为中间转换提供有限的奖励信号,尤其是在长视距操作任务中。一种自然的方法是建模中间任务进度,并将其用作策略改进的密集反馈。尽管现有进度感知方法在架构上存在差异,但它们通常将任务进度形式化为显式标量预测,这为建模中间观测与任务目标之间的关系提供的结构有限,可能阻碍有效的转换级信用分配。我们引入了进度场强化学习(PF-RL),该方法在预训练的VLA特征之上学习一种结构化的目标条件进度表示,并将其转换为用于策略优化的密集信用。一个轻量级的共享进度场头将当前表示和目标表示映射到一个紧凑的进度空间,其中几何距离诱导出目标条件价值,而互补的时间与目标结构目标则塑造所学习的几何结构。转换级价值变化自然产生密集的进度优势,从而能够为离线策略改进和在线强化微调提供细粒度的信用分配。在LIBERO、RoboTwin2.0以及真实世界双臂操作任务上的大量实验表明,PF-RL在强监督微调、强化微调和进度感知基线之上持续提升了策略性能。
英文摘要
Reinforcement Fine-Tuning~(RFT) has emerged as a promising paradigm for improving Vision-Language-Action~(VLA) policies, yet sparse task-level outcomes provide limited credit for intermediate transitions, especially in long-horizon manipulation. A natural approach is to model intermediate task progress and use it as dense feedback for policy improvement. Despite their architectural differences, existing progress-aware methods commonly formulate task progress as an explicit scalar prediction, providing limited structure for modeling how intermediate observations relate to the task goal, which may hinder effective transition-level credit assignment. We introduce Progress Field Reinforcement Learning (PF-RL), which learns a structured goal-conditioned progress representation over pretrained VLA features and converts it into dense credit for policy optimization. A lightweight shared Progress Field head maps current and goal representations into a compact progress space, where geometric distance induces goal-conditioned value, while complementary temporal and goal-structure objectives shape the learned geometry. Transition-level value changes naturally yield dense progress advantages, enabling fine-grained credit assignment for both offline policy improvement and online reinforcement fine-tuning. Extensive experiments on LIBERO, RoboTwin2.0, and real-world bimanual manipulation tasks show that PF-RL consistently improves policy performance over strong supervised fine-tuning, reinforcement fine-tuning, and progress-aware baselines.