发表机构
East China Normal University; Peking University; University of Science and Technology of China(华东师范大学; 北京大学; 中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MATE通过检索后处理将成功轨迹转化为面向执行的内存,提升具身智能体任务成功率并减少令牌消耗。
AI 中文摘要
经验复用可以减少具身智能体中的重复探索,但先前成功的轨迹可能不适用于当前执行环境。现有内存系统主要优化构建和检索;因此,当检索到的经验包含不兼容的动作或不合适的结构层级时,语义相关性和历史成功性仍显不足。我们引入面向任务条件执行的内存自适应(MATE),这是一种确定性的检索后处理流程,将轨迹转换为面向执行的内存。MATE移除过时的控制上下文,提取条件-动作-效果转换,应用验证过的动作归一化,选择任务相关的表示,并在固定预算下序列化结果,无需额外的LLM推理。在134个ALFWorld任务中,MATE使用Qwen2.5-14B和72B分别达到81.3%和93.3%的任务成功率,同时使用的令牌约为原始轨迹所需的十分之一。对照比较表明,验证过的动作归一化是MATE恢复检索经验效用的主要机制,支持内存自适应作为检索与具身执行之间的一个独立阶段。
英文摘要
Experience reuse can reduce repeated exploration in embodied agents, but a trajectory that succeeded previously may be unsuitable for the current execution context. Existing memory systems pri marily optimize construction and retrieval; semantic relevance and historical success therefore remain insufficient when retrieved ex perience contains incompatible actions or an inappropriate level of structure. We introduce Memory Adaptation for Task-Conditioned Execution (MATE), a deterministic post-retrieval procedure that converts trajectories into execution-oriented memory. MATE re moves obsolete control context, extracts condition-action-effect transitions, applies verified action normalization, selects a task dependent representation, and serializes the result under a fixed budget without additional LLM inference. On 134 ALFWorld tasks, MATE achieves task success rates of 81.3% and 93.3% with Qwen2.5-14B and 72B while using approximately one-tenth of the tokens required by raw trajectories. Controlled comparisons show that verified action normalization is the principal mechanism by which MATE restores the utility of retrieved experience, support ing memory adaptation as a distinct stage between retrieval and embodied execution.