GESTO:面向动态场景推理的以人为中心的时空记忆
GESTO: Human-Centric Spatio-Temporal Memory for Reasoning in Dynamic Scenes
- KTH Royal Institute of Technology(瑞典皇家理工学院)
- University of Stuttgart(斯图加特大学)
- Max Planck Research School for Intelligent Systems(马克斯·普朗克智能系统研究所)
- Google Research(谷歌研究院)
- TU Munich(慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
该研究提出GESTO时空记忆模型,结合4D场景图与人机交互、目标事件的两级结构,在基准测试中表现优异,可支撑动态人类环境下的以活动为中心的时空推理。
中文摘要 AI 辅助
在人类环境中运行的机器人需要记忆,不仅要捕捉存在的物体及其位置,还要捕捉人们随时间对这些物体的使用方式,以及个体交互如何组成目标导向的活动。现有的4D场景图会保留物体和地点的历史,但忽略活动结构;而活动表示要么未基于持久的3D场景,要么依赖外部提供的事件边界和物体关联。我们提出GESTO(Grounded Event and Spatio-Temporal memOry,即基于事件与时空的记忆),这是一种时空记忆,它将持久的4D场景图与原子级人机交互、目标驱动事件的两级层次结构相结合。GESTO从RGB-D观测流中自动提取带时间戳的交互,将其与持久场景实体关联,把它们分组为事件,并利用事件上下文优化不确定的物体关联。一个感知关系的工具调用智能体会查询生成的记忆,以进行以活动为中心的时空推理。我们在现有基准的可复现文本、二元和时间类别,以及40个新的Space2Event和Event2Space查询上对GESTO进行评估。GESTO在标准类别上分别获得0.71、0.75和0.70的分数,接近由真实事件和物体关联提供的方法的性能,同时在移除这些输入时,其性能显著优于相同的推理框架。它在Space2Event和Event2Space查询上进一步获得0.73和0.75的分数。消融实验表明,层次事件结构和上下文感知的关联优化提供了互补的益处,支持基于活动的层次记忆,用于动态人类环境中的回溯推理。
英文摘要
Robots operating in human environments need memories that capture not only what objects exist and where, but also how people use them over time and how individual interactions compose into goal-directed activities. Existing 4D scene graphs preserve object and place histories but omit activity structure, whereas activity representations are either not grounded in persistent 3D scenes or rely on externally provided event boundaries and object associations. We present GESTO (Grounded Event and Spatio-Temporal memOry), a spatio-temporal memory that couples a persistent 4D scene graph with a two-level hierarchy of atomic human--object interactions and goal-driven events. From an RGB-D observation stream, GESTO automatically extracts timestamped interactions, grounds them to persistent scene entities, groups them into events, and uses event context to refine uncertain object associations. A relation-aware tool-calling agent queries the resulting memory for activity-centric spatio-temporal reasoning. We evaluate GESTO on the reproducible text, binary, and time categories of an existing benchmark, together with 40 new Space2Event and Event2Space queries. GESTO achieves scores of 0.71, 0.75, and 0.70 on the standard categories, approaching a method supplied with ground-truth event and object grounding, while substantially outperforming the same reasoning framework when these inputs are removed. It further achieves 0.73 and 0.75 on Space2Event and Event2Space queries. Ablations show that hierarchical event structure and context-aware grounding refinement provide complementary benefits, supporting activity-grounded hierarchical memory for retrospective reasoning in dynamic human environments.