arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09410cs.RO

权重中的技能,代码中的记忆:面向依赖记忆的机器人操作的混合学习

Skills in Weights, Memory in Code: Hybrid Learning for Memory-Dependent Robot Manipulation

Yunhao Zhao, Zhenyang Ni, Haoyang Chen, Ruohan Zhang, Qi Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

针对现实机器人操作的非马尔可夫性,提出混合学习框架 HyMeS,结合编码智能体与马尔可夫 VLA,在 RoboMemArena 上显著提升操作成功率,实现数据高效的组合泛化。

中文摘要 AI 辅助

现代视觉-语言-动作(VLA)策略已掌握了广泛的操作技能,但通常仅从当前观测或短固定长度历史中生成每个动作块。然而,现实世界的操作往往是非马尔可夫的,要求机器人保留并推理长 horizon 交互历史中与任务相关的信息,以确定下一个动作。为应对这一挑战,我们提出 HyMeS,一个混合学习框架,利用编码智能体的推理和记忆管理能力来引导马尔可夫 VLA 完成依赖记忆的操作。具体而言,HyMeS 通过基于梯度的模仿学习学习低级运动技能,同时编码智能体通过启发式学习获取高级记忆管理策略,方法是从 rollout 反馈中迭代更新可执行的启发式系统。此外,我们通过多模态阶段完成验证闭合引导与执行的循环,该验证利用本体感受信号和多帧 VLM 判断更新记忆。与端到端的记忆增强型 VLA 相比,HyMeS 仅需要可复用运动技能的演示,而非每个依赖历史的任务配置,实现了数据高效的组合泛化。在 RoboMemArena 上,HyMeS 相较于 pi0.5 将平均累计成功率从 52.5% 提升至 66.2%,平均任务成功率从 41.3% 提升至 60.1%,同时在累计成功率上优于 PrediMem 4.5 个百分点,在任务成功率上优于 14.5 个百分点。

英文摘要

Modern vision-language-action (VLA) policies have acquired broad manipulation skills, but typically generate each action chunk from the current observation or a short fixed-length history. However, real-world manipulation is often non-Markovian, requiring robots to retain and reason over task-relevant information from long-horizon interaction histories to determine the next action. To address this challenge, we propose HyMeS, a hybrid learning framework that leverages the reasoning and memory-management capabilities of coding agents to steer a Markovian VLA for memory-dependent manipulation. Specifically, HyMeS learns low-level motor skills through gradient-based imitation learning, while a coding agent acquires high-level memory-management strategies through heuristic learning by iteratively updating an executable heuristic system from rollout feedback. Furthermore, we close the loop between steering and execution through multimodal stage-completion verification, which updates memory using proprioceptive signals and multi-frame VLM judgments. Compared with end-to-end memory-augmented VLAs, HyMeS requires demonstrations only for reusable motor skills rather than for every history-dependent task configuration, enabling data-efficient compositional generalization. On RoboMemArena, HyMeS improves mean cumulative success from 52.5% to 66.2% and mean task success from 41.3% to 60.1% over pi0.5, while outperforming PrediMem by 4.5 points in cumulative success and 14.5 points in task success.

补充信息

↑