AI 中文总结
本文提出Mimir,一种分离世界与任务记忆并具备动态grounding的神经符号记忆系统,在EB-ALFRED、EB-Habitat等具身任务中显著提升了智能体的长程执行成功率。
AI 中文摘要
长程具身任务要求智能体在部分可观测条件下行动,同时保留场景信念与执行进度。扁平历史或隐式策略状态可能包含过往观测,但未提供明确接口以判断哪些世界事实支持当前活跃目标。本文提出Mimir,一种神经符号记忆系统,将世界记忆与任务记忆分离,并在每次行动前对二者进行动态grounding。世界记忆存储物体位置、物体状态及感知证据;任务记忆存储有序目标议程、进度状态、手部状态、失败情况及执行约束。grounding模块将活跃目标与 recalled 世界候选绑定,填补缺失的源位置,并在规划及具身特定执行前附加证据。在测试的各骨干网络上,Mimir在不同EB-ALFRED与EB-Habitat任务中均有提升,最大增益分别为42.5%与23.0%。与相同骨干网络下评估的现有最优智能体及记忆系统结果相比,Mimir的整体平均成功率提升8.5%;在EB-Habitat长程子集上,Mimir达到86.0%的成功率,大幅优于当前闭源模型。我们的代码将很快发布。
英文摘要
Long-horizon embodied task requires agents to act under partial observability while preserving both scene belief and execution progress. Flat histories or implicit policy states may contain past observations, but they do not provide an explicit interface for deciding which world facts support the currently active goal. We introduce Mimir, a neuro-symbolic memory that separates world memory from task memory and dynamically grounds them before each action. World memory maintains object locations, object states, and perceptual evidence, while task memory maintains an ordered goal agenda, progress state, hand state, failures, and execution constraints. A grounding module binds the active goal to recalled world candidates, fills missing source locations, and attaches evidence before planning and embodiment-specific execution. Across tested backbones, Mimir consistently improves on different EB-ALFRED and EB-Habitat tasks, with maximum gains of 42.5% and average gains of 23.0%, respectively. Compared with the best results among prior agent and memory systems evaluated under the same backbone, Mimir improves the overall average success rate by 8.5%. Finally, on the EB-Habitat Long-horizon subset, Mimir achieves 86.0% success rate, substantially outperforming current closed-source models. Our code will be released soon.
Comments9 pages, 4 figures