发表机构
University of Maryland, College Park; The University of Texas at Austin; University of Illinois Urbana-Champaign(马里兰大学帕克分校; 德克萨斯大学奥斯汀分校; 伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究序列式具身问答中不同内存架构表现,发现仅保留现有记忆不足,短视情节数据训练的智能体有时间不匹配问题。强调结构化、基于空间记忆的必要性,实验表明其能打破准确性-效率权衡,对真实机器人连续智能操作至关重要。
AI 中文摘要
具身问答(EQA)传统上是在情节性框架下进行评估的,即智能体独立解决每个任务并在情节之间重置内部状态。然而,现实世界中的机器人是持续运行的,必须积累、保留并选择性地重用从先前交互中获取的信息。尽管有此实际需求,但在EQA中支持序列记忆所需的架构机制仍未得到充分探索。在这项工作中,我们研究了在对EQA智能体进行序列评估时,即在同一场景中回答多个问题且记忆在查询之间传递时,不同内存架构的表现。我们发现仅仅保留现有记忆往往是不够的。仅保留可遍历性信息(如二维占用地图)的智能体,记住了机器人探索过的位置,但没有记住后续问题所需的视觉语义证据。在短视情节数据上训练的智能体面临不同的挑战:当面对连续的多查询历史时,它们继承的上下文存在严重的时间不匹配,而不是形成可重用的场景表示。为了克服这一架构瓶颈,我们强调了结构化的、基于空间的记忆的必要性:将持久视觉观察映射到度量三维几何上的架构,在连贯的场景表示中保留视觉语义证据。在模拟环境中的大量实验表明,这种形式的记忆打破了序列设置中的准确性-效率权衡,同时实现了更高的答案准确性和更低的导航成本。我们还在真实世界的移动机器人上验证了这些发现,证明基于空间的视觉记忆对于在物理环境中实现连续、智能操作至关重要。
英文摘要
Embodied question answering (EQA) is traditionally evaluated under an episodic formulation, where agents solve each task independently and reset internal state between episodes. However, real-world robots operate continuously and must accumulate, retain, and selectively reuse information acquired from prior interactions. Despite this practical requirement, the architectural mechanisms needed to support sequential memory in EQA remain underexplored. In this work, we investigate how different memory architectures behave when EQA agents are evaluated sequentially, with multiple questions answered in the same scene while memory is carried forward across queries. We find that simply preserving existing memory is often insufficient. Agents that retain only traversability information, such as 2D occupancy maps, remember where the robot has explored but not the visual-semantic evidence needed for later questions. Agents trained on short-horizon episodic data face a different challenge: when exposed to continuous, multi-query histories, their inherited context suffers from severe temporal mismatch, rather than forming a reusable scene representation. To overcome this architectural bottleneck, we highlight the necessity of structured, spatially grounded memory: architectures that map persistent visual observations onto metric 3D geometry preserve visual-semantic evidence in a coherent scene representation. Extensive experiments in simulated environments reveal that this form of memory breaks the accuracy-efficiency tradeoff in sequential settings, simultaneously achieving higher answer accuracy and lower navigation costs. We further validate these findings on a real-world mobile robot, demonstrating that spatially grounded visual memory is critical for enabling continuous, intelligent operation in physical environments.
CommentsAccepted to IROS 2026