永不回望:从自我中心视频理解3D物体记忆中的持久性
Never Look Back: Understanding Persistence in 3D Object Memory from Egocentric Videos
浏览论文内容
中文总结 AI 辅助
提出Ledger持久3D物体记忆系统,从自我中心视频中结合位置、历史与描述,提升HD-EPIC和UCS-Bench准确率,并实现低误差的Ego4D物体定位。
中文摘要 AI 辅助
当我们在世界中移动并执行日常任务时,我们会遇到那些可能只在之后才变得相关的物体。我们能够回忆起把某物放在哪里,或者容器内有什么,即使当时并不知道以后会需要它。在此,我们研究一个具身助手如何通过观察一个人的日常活动,从自我中心视频中构建类似的记忆。我们提出了Ledger,一种持久的3D物体记忆,它结合了物体位置、其历史记录和上下文描述。它关联整个录制过程中的观察结果,并在物体离开视野后仍保留它们,包括那些人从未触碰过的物体。它通过静止位置对每个物体的观察结果进行聚类,并且仅在重复证据出现后才记录移动,从而减少定位噪声的影响。简短的描述保留了诸如物体内容或支撑表面等细节。它保存这些记录,以便日后无需访问原始图像或视频即可回答空间问题。我们的记忆将HD-EPIC的准确率从29.7%提升至42.6%,将UCS-Bench的准确率从33.8%提升至38.5%,并以0.99米的中位误差对Ego4D物体进行定位。我们的分析识别了时间持久性、上下文描述和检索的互补作用。我们对100个拼接流(每个包含多个场景)的研究进一步揭示了检索和构建中的失败。与单场景流相比,逐场景构建部分恢复了跨场景变化所损失的性能。
英文摘要
As we move through the world and carry out everyday tasks, we encounter objects that may become relevant only later. We are capable of recalling where we left something or what was inside a container, even without knowing we would need it later. Here, we study how an embodied assistant can build a similar memory from egocentric videos, by observing a person's day-to-day activities. We present Ledger, a persistent 3D object memory that combines object locations, their histories, and contextual descriptions. It associates observations across the recording and retains objects after they leave the view, including those the person never touches. It clusters each object's observations by resting locations and records a move only after repeated evidence, reducing the effect of localization noise. Short descriptions preserve details such as an object's contents or supporting surface. It saves these records to later answer spatial questions without having to access the original images or video. Our memory raises HD-EPIC accuracy from 29.7% to 42.6%, UCS-Bench accuracy from 33.8% to 38.5% and localizes Ego4D objects with a 0.99 m median error on returned predictions. Our analyses identify complementary roles for temporal persistence, contextual descriptions, and retrieval. Our study on 100 stitched streams of multiple scenes each further exposes failures in both retrieval and construction. Per-scene construction partially recovers the performance lost across scene changes compared to that of single scene streams.