arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PathTime-VLA:用于视觉-语言-动作策略因式化后训练的路径-时间解耦

PathTime-VLA: Path-Time Decoupling for Factorized Post-Training of Vision-Language-Action Policies

Qing Huang, Yifei Yang, Ziqing Zou, Anzhe Chen, Zhenjie Zhu, Yufei Wei, Rong Xiong, Yue Wang

arXiv 2610.11771首次发表:更新:

发表机构

Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出PathTime-VLA,通过路径-时间解耦优化VLA策略,经分阶段后训练后,在三项任务中成功率优于基线,且成功试验平均完成时间缩短39%-52%。

AI 中文摘要

视觉-语言-动作(VLA)策略通常以固定时间间隔预测动作,将机器人的运动路径与其执行节奏耦合在一起。这种耦合给遥操作的适配带来了困难:有用的几何指导会附带由界面延迟和操作者行为决定的时间特性。我们的核心见解是将经典运动规划的路径-时间参数化引入VLA的学习动作表示中。我们提出PathTime-VLA,它将运动表示为进度索引的交互路径$X(s)$和正区间时间轮廓。后者定义了一个单调时间律$t(s)$,生成控制器命令$X(s(t))$。对于给定路径,可通过时间轮廓表达替代执行方式,从而在不改变几何预测目标的情况下实现分块速度选择。该表示支持分阶段后训练流程:演示和DAgger干预建立目标领域先验,Speed-DQN从机器人交互中学习执行乘数,Path-AWR使用回滚结果优化扩散路径生成器。路径条件动作专家实现生成的运动,同时保持路径生成与执行计时的不同学习接口。在三项任务中,完整方法的成功率为58/60,而PathTime-VLA在固定1×下采用BC+DAgger的成功率为57/60,成功试验的平均完成时间缩短了约39%-52%。

英文摘要

Vision-Language-Action (VLA) policies typically predict actions at fixed time intervals, coupling the route a robot follows with its execution pace. This coupling complicates adaptation from teleoperation: useful geometric guidance comes with timing shaped by interface delays and operator behavior. Our key insight is to bring the path-time parameterization of classical motion planning into the learned action representation of a VLA. We introduce PathTime-VLA, which represents motion as a progress-indexed interaction path $X(s)$ and a positive interval-time profile. The latter defines a monotone time law $t(s)$, yielding controller commands $X(s(t))$. For a given path, alternative executions are expressed through the time profile, allowing chunk-wise speed choices without changing the geometric prediction target. This representation supports a staged post-training procedure: demonstrations and DAgger interventions establish a target-domain prior, Speed-DQN learns execution multipliers from robot interaction, and Path-AWR uses rollout outcomes to refine the diffusion path generator. A path-conditioned action expert realizes the resulting motions while maintaining distinct learning interfaces for path generation and execution timing. Across three tasks, the complete method achieves $58/60$ successes versus $57/60$ for PathTime-VLA under BC + DAgger at fixed $1\times$, with approximately $39$-$52\%$ shorter mean completion times over successful trials.

Comments8 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑