发表机构
Zhejiang University; The Hong Kong University of Science and Technology; The Hong Kong University of Science and Technology (Guangzhou); vivo Robotics Lab; South China Normal University(浙江大学; 香港科技大学; 香港科技大学(广州); vivo机器人实验室; 华南师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出轻量级插件模块PMTRM,通过编码有限历史并结合时间异质性等损失解决具身策略学习中的阶段模糊问题,在仿真与真实机器人实验中提升了含阶段模糊任务的成功率。
AI 中文摘要
机器人操纵通常包含重复动作,在不同阶段的局部观测看起来相似。当这些阶段需要不同动作时,主要依赖当前观测的策略可能会重复已完成的动作或在错误的时间切换阶段。为解决这种阶段模糊问题,我们提出伪记忆时间重编码模块(PMTRM),这是一个仅含761万参数的轻量级插件模块,可将执行状态和动作的有限历史编码为现有策略的潜在序列。为帮助区分阶段,时间异质性目标会惩罚该序列中远距离位置间的正相似度,而锚点损失和重构损失则保留动作预测所需的信息。重构解码器仅在训练期间使用,推理时由时间重编码器向策略提供历史信息。我们先在合成序列和机器人数据上逐步训练该模块,再结合策略进行联合训练,使用时间掩码适配部分历史。这种集成保留了原始动作头和动作空间,在原始策略损失中添加了辅助损失。在仿真环境和真实机器人上,针对多个策略骨干网络的实验表明,该模块在存在阶段模糊的任务上提升了任务成功率,且仅增加少量计算量。
英文摘要
Robotic manipulation often contains repeated motions whose local observations look similar at different phases. When these phases require different actions, a policy that relies mainly on the current observation may repeat completed motions or switch phases at the wrong time. To address this phase ambiguity, we present the Pseudo-Memory Temporal Re-encoding Module (PMTRM), a lightweight plug-in module with only 7.61M parameters that encodes a bounded history of executed states and actions into a latent sequence for existing policies. To help distinguish phases, a temporal heterogeneity objective penalizes positive similarity between distant positions in this sequence, while anchor and reconstruction losses preserve information needed for action prediction. The reconstruction decoder is used only during training, leaving the temporal re-encoder to supply history to the policy at inference. We train the module progressively on synthetic sequences and robot data, then jointly with the policy, using temporal masking to accommodate partial histories. This integration retains the original action head and action space and adds auxiliary losses to the original policy loss. Experiments with multiple policy backbones in simulation and on a real robot show improved task success on tasks with phase ambiguity, with little additional computation.