HODAgent:面向物理世界人机交互的按需、响应式人形机器人
HODAgent: Towards On-Demand, Responsive Humanoids for Physical World Human Interaction
浏览论文内容
中文总结 AI 辅助
HODAgent作为面向服务人形机器人的System-2具身智能体,通过半双工架构实现自适应服务,在仿真与实体机器人上均显著优于基准模型。
中文摘要 AI 辅助
我们提出了HODAgent,这是一种面向服务场景中人形机器人的System-2具身智能体,用于解决情境意图、响应式执行、任务修订和结果验证问题。其半双工架构整合了Env-Interactor、Planner、Executor和分层Memory,以在服务周期内维持连贯的交互、规划和任务状态,从而能够在运动过程中处理新请求、保留进度、修订动作,并基于执行结果确认任务完成。一个共享接口连接了仿真环境与实体机器人(Unitree G1),实现了平台特定控制的隔离。在包含164个案例的交互仿真中,HODAgent在两个VLM骨干模型下分别达到84.8%和91.5%的联合成功率,较基准模型分别提升9.8和18.9个百分点;在实体机器人上,其任务通过率分别为原子任务92%、复合任务72%、完整任务63.3%,在多个具身基准测试中较基准模型提升0.7至9.0个百分点。结果表明,统一的System-2智能体可实现跨仿真与现实的自适应人形机器人服务。
英文摘要
We propose HODAgent, a System-2 embodied agent for humanoid robots in service settings, addressing situated intent, responsive execution, task revision, and outcome verification. Its semi-duplex architecture integrates an Env-Interactor, Planner, Executor, and hierarchical Memory to maintain coherent interaction, planning, and task state during service episodes. This allows handling new requests during motion, retaining progress, revising actions, and grounding closure in execution outcomes. A shared interface connects simulation and physical robots (Unitree G1), isolating platform-specific control. In an interactive simulation with 164 cases, HODAgent achieves 84.8% and 91.5% Joint Success under two VLM backbones, outperforming baselines by 9.8 and 18.9 points. On physical robots, pass rates are 92% (atomic), 72% (composite), and 63.3% (complete tasks). On multiple embodied benchmarks, it improves over baselines by 0.7-9.0 points. Results show a unified System-2 agent enables adaptive humanoid service across simulation and reality.
发表机构
- Xiaopeng(小鹏)
机构由 AI 辅助整理,请以论文原文为准。