发表机构
University of Maryland, College Park(马里兰大学帕克分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对基础模型在具身导航部署中的训练偏差和上下文长度限制,提出TAP和MemCtrl方法,分别提升个性化目标寻找18%和长任务性能20%。
AI 中文摘要
我们提出并解决与在具身智能体执行导航任务时部署基础模型(FMs)相关的两个问题:1)基础模型中的训练偏差导致在未见环境中个性化能力差;2)基础模型有限的上下文长度阻碍了任务成功,尤其是在长时程任务中。针对前者,我们的解决方案涉及利用从场景中挖掘的人类习惯数据对基础模型进行预提示;针对后者,我们通过一种新颖的“记忆头”增强实现主动记忆管理。我们首先对基于基础模型的具身导航现有文献进行了分类,并强调了这些局限性。随后,我们提出了我们的方法,即交通感知规划(TAP)和MemCtrl,以解决这些局限性。借助TAP,我们在实验室环境中使用Turtlebot进行了真实世界实验,用于个性化目标寻找,结果显示相比非TAP基线平均提升了18%。关于MemCtrl,我们报告了在各种具身任务中平均提升6%,在长指令子集上提升20%,同时使用的上下文量仅为基线模型的一半。受这些结果的启发,我们阐述了关于基于基础模型的具身智能体在真实世界环境中可部署性的立场,并指出了开放的研究方向。
英文摘要
We present and tackle two problems associated with deploying Foundation Models (FMs) on Embodied Agents performing navigation: 1) Training bias in FMs leading to poor personalization in unseen environments, and 2) Limited FM context length hindering success, especially on long horizon tasks. Our solution for the former involves priming the FM with human-habit data mined from the scene and our solution for the latter involves active memory management via a novel `memory head' augmentation. We first present a taxonomy of existing literature on FM-based Embodied Navigation, and highlight these limitations. We then present our approaches, Transit-Aware Planning (TAP) and MemCtrl to address the limitations. With TAP, we present real-world results in a lab environment with a Turtlebot for personalized target finding that shows an average improvement of 18% over a non-TAP baseline. On MemCtrl, we report a 6% average improvement across various embodied tasks, with 20% on long instruction subsets, all while using nearly half the context used in the baseline model. Motivated by these result, we present our stance the deployability of FM-based embodied agents in real-world environments, and highlight open research directions.