arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PACE:面向长周期具身操纵的阶段进度感知信用分配框架

PACE: Phase-Progress-Aware Credit for Long-Horizon Embodied Manipulation

Chengye Song, Jiawei Zhang, Rui Song, Shengqi Wang, Xiangrong Zhang, Ziyi Wang, Huanbin Zhou, Hongzhou Wang

arXiv 2608.15026首次发表:更新:

AI 中文总结

PACE是面向长周期具身操纵的信用分配框架,通过GLC-Critic实现步骤级信用分配,结合PPD训练策略,在仿真与真实机械臂实验中较基线取得显著性能提升。

AI 中文摘要

视觉-语言-动作(VLA)模型的训练后优化通常依赖专家演示和策略交互轨迹。然而在长周期操纵任务中,单个回合往往包含数百个控制步骤和多个阶段,且任务的成功或失败仅在回合终止时才能判定。因此策略改进需要基于步骤级别的信用信号,以区分推动任务进展的行为与停滞或倒退的行为。本文提出PACE,一种针对长周期操纵任务训练后优化的信用分配框架,核心是阶段进度感知评论家。PACE包含两个关键模块:(1)全局-局部协作价值修正评论家(GLC-Critic),该模块聚合局部时间窗口内的视觉和运动差异特征,以推断每一步的阶段和阶段内进度,并据此对离散化的剩余代价分布应用残差修正,从而实现步骤级信用分配;(2)渐进式策略蒸馏(PPD),该模块通过任务相关阈值将信用转换为正负条件,并训练信用条件下的动作生成策略:首先用高信用正样本保护预训练策略,随后结合所有正负信用学习质量边界,推理阶段则通过条件输出的差异放大高信用行为。大量仿真实验和多样化的真实世界机械臂实验表明,PACE相较于最强基线方法始终取得显著提升。

英文摘要

Post-training of vision-language-action (VLA) models typically relies on expert demonstrations and policy interaction trajectories. However, in long-horizon manipulation, a single episode often spans hundreds of control steps and multiple phases, while success or failure is only revealed at episode termination. Policy improvement therefore requires step-level credit signals to distinguish behaviors that advance the task from those that stall or regress. We present PACE, a credit-assignment framework for post-training on long-horizon manipulation, centered on a phase-progress-aware critic. PACE consists of two key modules: (1) the Global-Local Cooperative Value-Correction Critic (GLC-Critic) aggregates visual and motion-difference features within local temporal windows to infer the phase and intra-phase progress of each step, and applies residual correction to a discretized remaining-cost distribution accordingly, enabling step-level credit assignment; (2) Progressive Policy Distillation (PPD) converts credit into positive and negative conditions via task-wise thresholds and trains a credit-conditioned action generation policy: it first protects the pretrained policy with high-credit positive samples, then incorporates all positive and negative credits to learn the quality boundary, and at inference amplifies high-credit behaviors through the difference between conditional outputs. Extensive simulation experiments and diverse real-world robotic-arm experiments demonstrate that PACE consistently achieves significant improvements over the strongest baseline.

Comments9 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑