发表机构
Samsung R&D Institute China - Beijing; Shanghai Jiao Tong University(三星中国研发研究院(北京); 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出R4DSG,一种面向长时序自我中心视频的相对4D场景图记忆,在EgoLifeQA数据集的对象相关问答任务中,较EgoRAG-Text取得显著性能提升,可作为可穿戴助手等的实用记忆载体。
AI 中文摘要
长时序自我中心视频是可穿戴AI助手的丰富信息载体,但诸如物品被移至何处、最后一次状态变更的时间或被重新放置的原因等以对象为中心的问题仍难以回答,因为基于字幕和文本的记忆很少能保留持久的对象身份或结构化的空间变化。现有的长视频问答方法主要侧重时间定位和片段检索,而现有的3D场景图方法通常假设输入比可穿戴RGB自由运动视频提供更强的几何信息,包括点云、RGB-D输入、带姿态的视图、稀疏重建或重建场景。R4DSG为长时序自我中心视频引入了一种相对4D场景图记忆。R4DSG不存储原始图序列,而是将视频转换为紧凑的可查询记忆条目,这些条目按时间、位置、持久对象、锚点相对变化和局部交互上下文进行索引。其核心思路是将稳定锚点与动态对象分离,在帧间维持持久的对象身份,并通过锚点相对转换来表示对象状态,而非采用全局对齐的世界模型。该方法基于近期仅RGB输入的可提示视频分割、时间传播和相对3D提升技术进展构建,生成可直接用于长时序问答的、支持检索的记忆。在EgoLifeQA中255个与对象相关的子集上的评估显示,在仅问题检索设置下,其整体性能较EgoRAG-Text提升6.7个百分点,在“何时”类问题上提升12.5个百分点,凸显了时间组织的对象记忆的价值。这些结果表明,相对4D场景图是可穿戴助手、AR系统和具身多媒体智能体的实用记忆载体。GitHub页面:this https URL。
英文摘要
Long-horizon egocentric video is a rich substrate for wearable AI assistants, but object-centric questions such as where an item was moved, when it last changed state, or why it was relocated remain difficult because caption- and transcript-based memories rarely preserve persistent object identity or structured spatial change. Existing long-video QA methods mainly emphasize temporal grounding and clip retrieval, while prior 3D scene-graph methods typically assume stronger geometry than free-motion wearable RGB video provides, including point clouds, RGB-D input, posed views, sparse reconstruction, or reconstructed scenes. R4DSG introduces a relative 4D scene graph memory for long egocentric video. Instead of storing raw graph sequences, R4DSG converts video into compact queryable memory entries indexed by time, place, persistent objects, anchor-relative change, and local interaction context. The main idea is to separate stable anchors from dynamic objects, maintain persistent object identity across frames, and represent object state through anchor-relative transitions rather than a globally aligned world model. Built on recent RGB-only advances in promptable video segmentation, temporal propagation, and relative 3D lifting, the method produces a retrieval-ready memory directly usable for long-horizon question answering. Evaluation on a 255-question object-related subset from EgoLifeQA shows, under question-only retrieval, a 6.7-point overall gain over EgoRAG-Text and a 12.5-point gain on when questions, which highlights the value of temporally organized object memory. These results position relative 4D scene graphs as a practical memory substrate for wearable assistants, AR systems, and embodied multimedia agents. GitHub Page: https://dualtransparency.github.io/R4DSG/.
Comments10 pages, 3 figures, ACM Multimedia 2026, egocentric video; 3D scene graph; temporal memory; graph retrieval; object-state reasoning; multimodal question answering