arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SplineWAM:基于B样条表示的世界动作模型自适应动作时域

SplineWAM: Adaptive Action Horizons for World Action Models via B-Spline Representations

Jun Guo, Xiaoshen Han, Qiwei Li, Nan Sun, Peiyan Li, Heyun Wang, Hang Lai, Weinan Zhang, Xinghang Li, Huaping Liu

arXiv 2609.39873首次发表:更新:

发表机构

Tsinghua University; Xiaomi Robotics; Shanghai Jiao Tong University; Peking University; CASIA(清华大学; 小米机器人; 上海交通大学; 北京大学; 中国科学院自动化研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SplineWAM利用B样条表示自适应压缩动作轨迹,动态调整执行时域,提高世界动作模型的吞吐量和成功率,在多个基准和真实机器人任务上优于基线。

AI 中文摘要

世界动作模型(WAMs)是大型具身策略,它们联合预测未来视频和待执行的动作,每次推理调用输出固定长度的动作块。这种策略在时间上均匀分配其计算预算,无法在自由空间运动中执行更长时间,也无法在接触丰富的操作中投入更多推理,这限制了WAM在云端服务时的吞吐量。我们提出SplineWAM,它将动作轨迹自适应地压缩为固定大小的三次B样条参数窗口,将节点时间拟合到运动特征上。一个参数预算随后解码为具有不同时间分辨率和时长的动作块,并且执行跨度以及到下一次策略调用的间隔均由预测本身决定。将视频监督对齐到演示的拟合节点时间而非均匀网格,将监督帧集中在动作轨迹复杂的区域。针对异步部署,我们引入了雅可比拉回实时分块(JP-RTC),该方法对机器人执行的解码原始动作施加块连续性,而非对样条参数施加,并通过解码器修正参数,使执行的前缀与已提交的动作一致。在LIBERO-Plus和RoboCasa上,SplineWAM相比动作分块WAM将成功率提高了8.2和4.4个百分点,同时将每个回合的策略调用次数减少了22%和26%。在三个双臂真实机器人任务中,在异步执行下,它领先或匹配基线,同时每次调用解码的执行运动量是基线的1.2到1.6倍。

英文摘要

World action models (WAMs) are large embodied policies that jointly predict future video and the actions to execute, emitting a fixed-length action chunk per inference call. Such a policy allocates its computational budget uniformly in time, unable to execute for longer over free-space motion or to spend more inference on contact-rich manipulation, which limits the throughput a WAM can reach when served in the cloud. We present SplineWAM, which adaptively compresses the action trajectory into a fixed-size window of cubic B-spline parameters, fitting the knot times to the characteristics of the motion. One parameter budget then decodes into chunks of varying temporal resolution and duration, and both the executed span and the interval until the next policy call follow from the prediction itself. Aligning the video supervision to the fitted knot times of the demonstration rather than to a uniform grid concentrates the supervised frames where the action trajectory is complex. For asynchronous deployment we introduce Jacobian-Pullback Real-Time Chunking (JP-RTC), which imposes chunk continuity on the decoded raw actions the robot executes rather than on the spline parameters, and corrects the parameters through the decoder so that the executed prefix agrees with the actions already committed. On LIBERO-Plus and RoboCasa, SplineWAM improves success rate over an action chunking WAM by $8.2$ and $4.4$ points while cutting policy calls per episode by 22% and 26%. On three bimanual real-robot tasks under asynchronous execution, it leads or matches the baseline while decoding 1.2 to 1.6 times as much executed motion per call.

CommentsProject website: https://splinewam.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑