ABot-AgentOS:一个具有终身多模态记忆的通用机器人智能体操作系统
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory
浏览论文内容
中文总结 AI 辅助
研究针对长期具身智能体运行需求,提出ABot-AgentOS通用机器人智能体操作系统,引入EmbodiedWorldBench基准,采用通用多模态图记忆及失败驱动自我进化循环,在相关测试中表现良好,证明该系统层可提升具身执行并提供持久记忆。
中文摘要 AI 辅助
近期的视觉语言模型(VLM)和视觉语言动作模型(VLA)系统提升了机器人感知和动作预测能力,但长期的具身智能体仍需要一个通用运行时层用于推理、记忆、工具使用、验证和跨实体执行。我们提出了ABot-AgentOS,一个位于低级控制器之上的通用机器人智能体操作系统,为场景条件规划、上下文隔离技能执行、多阶段验证、多模态记忆和边缘云协作提供审议智能体层。为评估此类系统,我们引入了EmbodiedWorldBench,一个具有16个室内、室外和混合场景、四个难度级别以及200多个涉及导航、对象搜索、NPC对话、动态事件和基于轨迹评分任务的可执行基准。ABot-AgentOS还引入了通用多模态图记忆,一个持久的源接地基板,将对话、视觉观察、空间上下文、时间关系和任务轨迹转换为类型化节点和边。一个失败驱动的自我进化循环将诊断出的记忆失败转换为门控运行时进化资产,仅提升到后续评估分割,防止当前分割的地面真值泄漏,同时实现持续改进。在初始的EmbodiedWorldBench子集中,ABot-AgentOS在任务成功和目标完成方面均优于单控制器基线。在记忆基准测试中,ABot-AgentOS静态版本在LoCoMo上达到87.5,在OpenEQA EM-EQA上达到59.9,在Mem-Gallery上达到88.6,在NExT-QA上Acc@All达到76.5;自我进化进一步将LoCoMo提升到88.7,OpenEQA提升到60.4,Mem-Gallery提升到89.0。这些结果表明,一个通用的智能体操作系统层可以改善长期的具身执行,同时为持续交互提供持久、可审计的记忆。
英文摘要
Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime layer for reasoning, memory, tool use, verification, and cross-embodiment execution. We present ABot-AgentOS, a general robotic Agent Operating System that sits above low-level controllers and provides a deliberative agent layer for scene-conditioned planning, context-isolated skill execution, multi-stage verification, multi-modal memory, and edge-cloud collaboration. To evaluate such systems, we introduce EmbodiedWorldBench, an executable benchmark with 16 indoor, outdoor, and hybrid scenes, four difficulty levels, and over 200 tasks involving navigation, object search, NPC dialogue, dynamic events, and trace-grounded scoring. ABot-AgentOS further introduces Universal Multi-modal Graph Memory, a persistent source-grounded substrate that converts dialogue, visual observations, spatial context, temporal relations, and task traces into typed nodes and edges. A failure-driven self-evolution loop converts diagnosed memory failures into gated runtime evo-assets that are promoted only to later evaluation splits, preventing current-split ground-truth leakage while enabling continual improvement. On an initial EmbodiedWorldBench subset, ABot-AgentOS improves over a single-controller baseline in both task success and goal completion. Across memory benchmarks, ABot-AgentOS Static achieves 87.5 on LoCoMo, 59.9 on OpenEQA EM-EQA, 88.6 on Mem-Gallery, and 76.5 Acc@All on NExT-QA; self-evolution further improves LoCoMo to 88.7, OpenEQA to 60.4, and Mem-Gallery to 89.0. These results suggest that a general Agent OS layer can improve long-horizon embodied execution while providing persistent, auditable memory for continual interaction.
发表机构
- AMAP CV Lab(AMAP计算机视觉实验室)
机构由 AI 辅助整理,请以论文原文为准。