发表机构
Motubrain Team(Motubrain团队)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对世界动作模型推理延迟导致机器人执行不连续的问题,通过10Hz双机械臂机器人对比6种异步部署策略,发现前缀条件生成可在任务性能、速度与平滑度间实现最佳平衡,为高延迟世界动作模型的实时部署提供了指导。
AI 中文摘要
世界动作模型通过迭代去噪生成固定时长的动作块,会产生显著的推理延迟,可能导致机器人执行过程中出现停顿、动作过时和不连续问题。本文对异步部署策略展开实证研究,该策略可将模型推理与动作执行重叠,以实现响应式、平滑的控制。我们在10Hz双机械臂机器人上对比了6种策略,包括同步执行、纯异步切换、事后动作融合、去噪时融合、推理时速度引导和前缀条件生成。评估结合了离线轨迹分析与在线实验,实验场景涵盖动态操作、精度要求高的放置任务及长时程任务。结果表明,观测、预测与执行指令间的准确时间对齐是基本要求,对齐误差会产生持续的动作块边界不连续问题,仅通过融合无法修正;对齐适当时,直接动作加权可作为简单平滑的基线,但在精度要求高的任务中会牺牲准确性;推理时速度引导在我们的平台上无法可靠约束已提交的动作;相比之下,前缀条件生成通过在训练中学习一致的动作延续,在任务性能、执行速度和轨迹平滑度之间实现了最佳整体平衡。这些发现明确了异步部署策略间的实际权衡,为在实时机器人系统中部署高延迟的世界动作模型提供了指导。
英文摘要
World Action Models generate fixed-horizon action chunks through iterative denoising, creating substantial inference latency that can cause pauses, stale actions, and discontinuities during robotic execution. We present an empirical study of asynchronous deployment strategies that overlap model inference with action execution to enable responsive and smooth control. We compare six strategies, including synchronous execution, pure asynchronous switching, post-hoc action blending, denoising-time blending, inference-time velocity guidance, and prefix-conditioned generation, on a 10 Hz bimanual robot. Evaluation combines offline trajectory analysis with online experiments across dynamic manipulation, precision-critical placement, and long-horizon tasks. Our results identify accurate temporal alignment between observations, predictions, and executed commands as a fundamental requirement. Alignment errors produce persistent chunk-boundary discontinuities that cannot be corrected through blending alone. With proper alignment, direct action weighting provides a simple and smooth baseline but sacrifices accuracy in precision-critical tasks. Inference-time velocity guidance fails to reliably constrain committed actions on our platform. In contrast, prefix-conditioned generation achieves the best overall balance between task performance, execution speed, and trajectory smoothness by learning consistent action continuations during training. These findings clarify the practical trade-offs among asynchronous deployment strategies and provide guidance for deploying high-latency World Action Models in real-time robotic systems.