发表机构
University of Georgia; Northeastern University; EmbodyX Inc.(佐治亚大学; 东北大学; EmbodyX公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对家庭机器人在执行任务时需整合新指令的问题,提出双层级框架CIRRA,结合LLM语义推理与结构整合,保留执行主干,减少冗余,在CHIRP基准上决策一致性达74.2%,并在人形机器人上验证有效性。
AI 中文摘要
家庭机器人必须在执行持续任务的同时容纳新的用户指令。现有智能体常常重新生成或大幅修订剩余任务序列,导致计划模糊、逻辑不一致和冗余执行。我们提出了持续指令协调问题,并提出了CIRRA(面向机器人智能体的持续指令协调),一个结合基于LLM的语义推理与规则约束的结构整合的双层级框架。CIRRA首先将传入的指令落地为独特的可执行技能,并解决未明确指定的动作和执行位置。然后,它保留正在进行的子任务序列作为执行主干,并通过将传入的子任务插入到位置匹配的片段中来生成整合候选。语义推理器仅评估修改后的片段,以识别依赖关系和冲突,并选择最逻辑连贯的局部整合。这种结构保持过程维持与持续执行的一致性,缓解模糊性和不一致性,并重用共享子任务以减少冗余执行。我们还引入了CHIRP(持续家庭指令协调与规划),一个基于文本的基准,包含八个家庭环境和六类日常活动的120个情节。在CHIRP上,CIRRA实现了74.2%的决策一致性,比最强的重新规划基线高出30个百分点;每个正确的融合决策都产生一个正确放置、无冲突的调度。在Unitree G1人形机器人上,CIRRA在每次试验中都在正确时刻中断正在进行的技能,并在所有指标上显著优于所有基线。
英文摘要
Household robots must accommodate new user instructions while executing ongoing tasks. Existing agents often regenerate or extensively revise the remaining task sequence, introducing plan ambiguity, logical inconsistency, and redundant execution. We formulate continual instruction reconciliation and propose CIRRA (Continual Instruction Reconciliation for Robot Agents), a dual-level framework combining LLM-based semantic reasoning with rule-constrained structural integration. CIRRA first grounds incoming instructions to unique executable skills and resolves underspecified actions and execution locations. It then preserves the ongoing subtask sequence as an execution backbone and generates integration candidates by inserting incoming subtasks into location-matched segments. The semantic reasoner evaluates only modified segments to identify dependencies and conflicts and select the most logically coherent local integration. This structure-preserving process maintains alignment with ongoing execution, mitigates ambiguity and inconsistency, and reuses shared subtasks to reduce redundant execution. We also introduce CHIRP (Continual Household Instruction Reconciliation and Planning), a text-based benchmark of 120 episodes across eight household environments and six categories of everyday activities. On CHIRP, CIRRA achieves 74.2% decision agreement, exceeding the strongest replanning baseline by 30 percentage points; every correct fusion decision yields a correctly placed, conflict-free schedule. On a Unitree G1 humanoid, CIRRA interrupts ongoing skills at the correct moment in every trial and significantly outperforms all baselines on every metric.