arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DriftOPD:用于单步VLA策略的序列级反向KL蒸馏

DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies

Youngjun Jun, Kyumin Choi, Youngmin Kim, Seonghyun Jin, Sunwoo Park, Jangho Park, Jong Chul Ye

arXiv 2610.00317首次发表:更新:

发表机构

KAIST; Sungkyunkwan University(韩国科学技术院; 成均馆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对VLA策略块级训练忽视长时程任务成功的问题,提出无需教师和展开的DriftOPD框架,通过分解序列级反向KL散度并优化单步漂移目标与Q函数,实现仅用离线数据的单步蒸馏,性能优于现有基线并媲美多步教师策略。

AI 中文摘要

视觉-语言-动作(VLA)模型日益依赖于在退缩时域控制下生成短动作块的动作专家。虽然块级训练在不同机器人形态上很方便,但它优化了局部动作似然,而没有明确考虑长时程任务成功。序列级强化学习可以解决这一局限,但通常需要策略展开和闭环交互,这对于真实机器人操作来说成本高昂。我们提出了DriftOPD,一个无需教师、无需展开的框架,用于连续VLA动作专家的序列级在策略蒸馏。我们表明,序列级反向Kullback-Leibler(KL)散度分解为块级反向KL项和一个未来势项,后者捕捉当前动作的长时程效应。DriftOPD分别使用单步漂移目标和从离线演示中学习的Q函数评论家来优化这两项,从而仅需离线数据和单步动作生成即可实现序列级优化。在仿真和真实世界操作中的多种VLA架构上,DriftOPD通常优于现有的单步蒸馏基线,同时达到与多步教师策略相当的任务成功率。这些结果表明,长时程行为可以有效地蒸馏到单步VLA动作专家中,而无需在线交互或单独的教师。

英文摘要

Vision-Language-Action (VLA) models increasingly rely on action experts that generate short action chunks under receding-horizon control. While chunk-level training is convenient across robot embodiments, it optimizes local action likelihood without explicitly accounting for long-horizon task success. Sequence-level reinforcement learning can address this limitation, but typically requires policy rollouts and closed-loop interaction, which are costly for real-robot manipulation. We introduce DriftOPD, a teacher-free, rollout-free framework for sequence-level on-policy distillation of continuous VLA action experts. We show that the sequence-level reverse Kullback-Leibler (KL) divergence decomposes into a chunk-level reverse-KL term and a future-potential term that captures the long-horizon effect of the current action. DriftOPD optimizes these two terms using a one-step drifting objective and a Q-function critic learned from offline demonstrations, respectively, enabling sequence-level optimization with only offline data and one-step action generation. Across multiple VLA architectures in simulation and real-world manipulation, DriftOPD generally outperforms existing one-step distillation baselines while achieving task success performance comparable to multi-step teacher policies. These results demonstrate that long-horizon behavior can be effectively distilled into one-step VLA action experts without online interaction or a separate teacher.

CommentsPreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑