发表机构
Princeton University; Arizona State University; Siemens Foundation Technologies(普林斯顿大学; 亚利桑那州立大学; 西门子基础技术研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出PHRetune离线方法,利用端口-哈密顿模型从演示中推导控制器增益,无需评估rollout或搜索,提升扩散及VLA策略在仿真和真实任务中的成功率并降低加速度。
AI 中文摘要
扩散策略和VLA策略在操作任务中通常通过下游阻抗控制器部署。这些控制器的刚度和阻尼增益影响任务成功率,但通常继承自数据采集过程,而非针对部署策略进行选择。尽管经验性增益扫描可以提升性能,但需要重复的评估 rollout。我们提出PHRetune,一种离线方法,无需评估 rollout 或增益搜索即可为冻结策略推导控制器增益。我们的方法从演示中学习端口-哈密顿模型,以估计策略预测动作相关的力和能量。将策略应用于记录的演示观测,并针对从演示推导的力和能量预算评估其预测。通过比较,我们以闭式形式推导出单一增益缩放因子,调整下游控制器同时保留策略及其动作表示。增益在评估前固定,无需先验机械臂模型、任务奖励、策略重训练或额外运行时计算。在LIBERO基准套件中,PHRetune将扩散策略成功率提升高达9.4个百分点,且推导的增益在经验性增益扫描中达到最高观测成功率。在所有四个真实世界操作任务中,PHRetuned扩散策略优于名义策略、替代增益调优方法和策略重训练基线。相同流程在SmolVLA和OpenVLA-OFT上提升每个任务的成功率,同时降低两种VLA骨干网络的加速度和加加速度。
英文摘要
Diffusion and VLA policies for manipulation are often deployed through downstream impedance controllers. The stiffness and damping gains of these controllers affect task success, yet are commonly inherited from data collection rather than selected for the deployed policy. Although empirical gain sweeps can improve performance, they require repeated evaluation rollouts. We introduce PHRetune, an offline method that derives controller gains for a frozen policy without evaluation rollouts or gain search. Our approach learns a port-Hamiltonian model from demonstrations to estimate the effort and energy associated with the policy's predicted actions. The policy is applied to recorded demonstration observations, and its predictions are assessed against demonstration-derived effort and energy budgets. From this comparison, we derive a single gain scale in closed form, adjusting the downstream controller while preserving the policy and its action representation. The gains are fixed before evaluation, without requiring a prior manipulator model, task rewards, policy retraining, or additional runtime computation. Across LIBERO suites, PHRetune improves Diffusion Policy success by up to 9.4 percentage points, with the derived gains achieving the highest observed success rates in empirical gain sweeps. On all four real-world manipulation tasks, PHRetuned Diffusion Policy outperforms the nominal policy, alternative gain-tuning methods, and a policy-retraining baseline. The same procedure improves success with SmolVLA and OpenVLA-OFT on every task, while reducing acceleration and jerk for both VLA backbones.