arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30378cs.ROcs.AI

PAVE:面向世界-动作策略的预测对齐与价值引导进化

PAVE: Predictive Alignment and Value-Guided Evolution for World-Action Policies

Botong Zhao, Fang Yu, Tim Yu, Senhua Zhu, Xinyuan Chen, Yue Lu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出PAVE策略,通过多时间步转换对齐与分布价值评论者优化,在三个模拟基准中实现最优性能,同时保留在线直接动作生成路径。

中文摘要 AI 辅助

直接的视觉-语言-动作策略可高效生成连续机器人动作,但标准行为克隆存在两个互补性缺陷:其表征未被明确要求描述场景如何在多个时间尺度上演化,且质量不等的部署轨迹常被重复使用,未区分有用动态与不良行为。我们提出PAVE(即本研究的\textit{\textbackslash method}),这是一种结合结果无关预测学习与结果感知策略改进的直接世界-动作策略。PAVE首先保留局部固定偏移JEPA目标,并在剩余episode的25%、50%、75%和100%处添加轨迹相对多时间步转换对齐。这些仅用于训练的目标要求当前策略表征既保留局部物理变化,又保留更长范围的任务进展,且无需向动作头提供显式未来token。随后,PAVE在累积部署轨迹上训练独立的分布价值评论者,计算与动作块对齐的N步优势,并将其转换为流匹配演员的正、负或空文本条件。因此,每条有效轨迹均可传授实际发生的情况,而演员仅在与相对更优动作相关的条件下部署。多时间步预测器和评论者在在线执行时被移除,仅保留从当前观测、语言指令和本体感觉生成直接动作。在三个模拟基准中,PAVE实现了最强的整体性能,同时保留了直接演员的在线执行路径。

英文摘要

Direct vision-language-action policies generate continuous robot actions efficiently, but standard behavior cloning leaves two complementary gaps: their representations are not explicitly required to describe how the scene evolves over multiple time scales, and deployment trajectories of unequal quality are often reused without separating useful dynamics from undesirable behavior. We introduce \method, a direct world-action policy that combines outcome-agnostic predictive learning with outcome-aware policy improvement. \method first retains a local fixed-offset JEPA objective and adds trajectory-relative multi-horizon transition alignment at 25%, 50%, 75%, and 100% of the remaining episode. These training-only targets require the current policy representation to preserve both local physical changes and longer-range task progress, without supplying explicit future tokens to the action head. \method then trains an independent distributional value critic on cumulative deployment trajectories, computes action-chunk-aligned $N$-step advantages, and converts them into positive, negative, or null text conditions for a flow-matching actor. Thus, every valid trajectory can teach what physically happened, while the actor is deployed only under the condition associated with relatively better actions. The multi-horizon predictor and critic are removed from online execution, preserving direct action generation from the current observation, language instruction, and proprioception. \redclaim{Across the three simulation benchmarks, \method achieves the strongest overall performance while preserving the direct actor's online execution path.}

发表机构

  • East China Normal University(华东师范大学)
  • EBKernel Shanghai Artificial Intelligence Laboratory(上海人工智能实验室EBKernel)

机构由 AI 辅助整理,请以论文原文为准。

↑