rMuscle:用于高效视觉-语言-动作模型推理的机器人肌肉记忆
rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference
浏览论文内容
中文总结 AI 辅助
针对VLA模型推理延迟问题,提出受肌肉记忆启发的rMuscle框架,利用跨执行相似性通过双阶段缓存减少计算和权重访问,在多种任务上实现1.29-1.42倍加速且保持成功率。
中文摘要 AI 辅助
工厂工作是具身智能一个有前景的早期场景:将重复性人工任务分配给机器人具有明确的经济回报,且结构化的工位使这些任务对当前策略而言易于处理。视觉-语言-动作(VLA)模型现在主导着这些机器人的策略范式。VLA模型的推理延迟直接影响机器人的响应速度和运动平滑性。然而,现有的VLA推理框架并未充分利用具身工作负载的特性,也未考虑VLA推理不同阶段的不同瓶颈。在本文中,我们首先刻画了具身工作负载,并识别出重复机器人执行过程中存在显著的任务相似性。我们进一步发现,这种相似性不仅存在于观测和动作轨迹中,还延伸到内部模型状态。基于这些观察,我们提出了rMuscle,一个受人类肌肉记忆启发的实时VLA推理框架。它通过双阶段肌肉记忆缓存利用跨执行相似性:上下文缓存重用视觉标记输出以减少计算,而动作缓存重用神经元激活模式以减少权重访问。我们通过在线缓存重计算、滑动窗口缓存检索以及连续去噪步骤间的掩码共享,将缓存内存占用和访问开销保持在较低水平。rMuscle在RTX 4090和Jetson Thor上,在LIBERO、RoboTwin和物理操作任务中实现了1.29-1.42倍的加速,同时在真实机器人上保持了原始成功率。
英文摘要
Factory work is a promising early scenario for embodied AI: assigning repetitive manual jobs to robots has clear economic payoff, and a structured station keeps the jobs tractable for current policies. Vision-Language-Action (VLA) models now dominate as the policy paradigm for these robots. The inference latency of VLA models directly affects robot responsiveness and motion smoothness. However, existing VLA inference frameworks do not fully exploit the characteristics of embodied workloads or account for the distinct bottlenecks across different stages of VLA inference. In this paper, we first characterize embodied workloads and identify substantial task similarity across repeated robot executions. We further find that such similarity extends beyond observations and action trajectories to internal model states. Drawing on these observations, we present rMuscle, a real-time VLA inference framework inspired by human muscle memory. It exploits cross-execution similarity through a dual-phase muscle-memory cache. The Context Cache reuses visual-token outputs to reduce computation, while the Action Cache reuses neuron activation patterns to reduce weight accesses. We keep both the cache memory footprint and access overhead low through online cache recomputation, sliding-window cache retrieval, and mask sharing across consecutive denoising steps. rMuscle achieves 1.29-1.42X speedup on RTX 4090 and Jetson Thor across LIBERO, RoboTwin, and physical manipulation tasks, while maintaining the original success rates on real-world robots.
发表机构
- Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University(上海交通大学并行与分布式系统研究所)
机构由 AI 辅助整理,请以论文原文为准。