发表机构
School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen); Jiangsu Cytoderm Intelligent Technology Co., Ltd.; National University of Singapore(哈尔滨工业大学(深圳)计算机科学与技术学院; 江苏赛克德智能科技有限公司; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对视觉-语言-动作模型动作块的弱结构化问题,提出埃尔米特轨迹先验,其中埃尔米特正则化变体在仿真与真实机器人任务中显著提升了操纵成功率且无额外推理开销。
AI 中文摘要
尽管近期面向机器人操纵的视觉-语言-动作(VLA)模型取得了进展,但动作块仍是一种弱结构化接口。现有工作通常将每个块展平为时间步控制量,依赖隐式数据学习,这在物理执行时表现为锯齿状运动和边界不连续性。为解决这些局限,我们引入埃尔米特轨迹先验,将块轨迹参数化为由端点位置和速度定义的分段三次埃尔米特曲线,以显式保证平滑性与连续性。我们通过三种变体在离散自回归和连续生成范式中实现了该固定算子:(1)埃尔米特令牌,自回归地预测量化边界变量;(2)埃尔米特支架,将干净动作分解为基础支架与残差;(3)埃尔米特正则化,严格将该先验作为辅助训练目标。在仿真基准和真实机器人平台上,埃尔米特正则化在这三种变体中取得了更优性能:在LIBERO上将π0.5基线成功率从95.9%提升至98.7%,在LIBERO-plus上从85.7%提升至90.9%,在四项真实机器人任务上从63.4%提升至90.0%,且无额外推理开销。轨迹分析表明,显式构建轨迹先验作为学习归纳偏置而非运行时约束最为有效。
英文摘要
Despite recent progress in Vision-Language-Action (VLA) models for robotic manipulation, the action chunk remains a weakly structured interface. Existing work typically flatten each chunk into per-timestep controls, relying on implicit data learning that manifests as jagged motion and boundary discontinuities during physical execution. To address these limitations, we introduce Hermite trajectory priors, parameterizing the chunk trajectory as a piecewise cubic Hermite curve defined by endpoint positions and velocities to explicitly enforce smoothness and continuity. We instantiate this fixed operator across discrete autoregressive and continuous generative paradigms via three variants: (1) Hermite Tokens, which predict quantized boundary variables autoregressively; (2) Hermite Scaffold, which decomposes clean actions into a base scaffold and residuals; and (3) Hermite Regularization, which applies the prior strictly as an auxiliary training objective. Across simulation benchmarks and real-robot platforms, Hermite Regularization achieves superior performance among these three variants, improving π0.5 baseline success rates from 95.9% to 98.7% on LIBERO, 85.7% to 90.9% on LIBERO-plus, and 63.4% to 90.0% across four real-robot tasks without additional inference overhead. Trajectory analyses reveal that explicitly structuring trajectory priors serves most effectively as a learning inductive bias rather than a runtime constraint.
CommentsProject page is available at https://aopolin-lv.github.io/Hermite/