arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ECHO:人类携带物体的具身相机观测

ECHO: Embodied Camera Observations of Human Object Carrying

Xuefei Sun, Lorin Achey, Kali Hamilton, Alberto Speranzon, Gregory Grebe, Yonatan Bisk, Christoffer Heckman

arXiv 2610.10438首次发表:更新:

发表机构

University of Colorado Boulder; Lockheed Martin; Carnegie Mellon University(科罗拉多大学博尔德分校; 洛克希德·马丁公司; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出上下文物体放置基准任务及ECHO合成数据集,通过RGB-D扫描与人类携带行为记录,验证联合推理场景、活动与上下文对预测物体目的地至关重要。

AI 中文摘要

具身智能体和辅助智能体不仅要识别物体,还必须根据环境布局以及居住者的习惯,推理出物体应放置的位置。该问题的研究进展一直受限,部分原因是缺乏专门的基准或数据集来定义和评估该任务。现有的RGB-D扫描数据集重建了没有人类活动的静态房间,而人-物交互数据集则捕捉了运动,但缺乏可导航的、完全重建的场景,也缺乏物体自然目的地的真值概念。我们提出了上下文物体放置这一基准任务:在观测到的物体携带过程中预测物体的目的地。为支持该任务,我们提出了人类携带物体的具身相机观测(ECHO),这是一个大规模合成数据集,将室内场景的密集RGB-D扫描与具身智能体携带日常物品到符合上下文的目的地的记录配对。ECHO是首个公开可用的数据集,结合了重建场景、人类活动、自然语言和上下文放置标注。它包含115个HM3D场景中159个楼层上的3,805个人工标注片段,涉及198个不同物体。每个楼层包含完整的RGB-D扫描,带有标注房间标签和表面列表。每个片段提供同步的RGB-D遭遇片段;6自由度相机、人和物体的轨迹;起始和目的表面;动作描述;以及一条人工撰写的上下文:一句话描述居住者的日常习惯,暗示目的地而不直接命名。我们使用输入掩蔽探针和端到端基线评估上下文物体放置。结果表明,没有任何单一输入模态是充分的,突显了联合推理场景结构、人类活动和上下文知识的必要性。

英文摘要

Embodied and assistive agents must do more than recognize objects: they must reason about where an object belongs given the layout of an environment and the habits of the people who live in it. Progress on this problem has been limited, in part because no dedicated benchmark or dataset exists to define and evaluate it. Existing RGB-D scan datasets reconstruct static rooms without human activity, while human-object-interaction datasets capture motion without a navigable, fully reconstructed scene or a ground-truth notion of an object's natural destination. We introduce contextual object placement as a benchmark task: predicting an object's destination during an observed object-carrying episode. To support this task, we present Embodied Camera observations of Human Object carrying (ECHO), a large-scale synthetic dataset that pairs dense RGB-D scans of indoor scenes with recordings of an embodied human carrying everyday objects to context-appropriate destinations. ECHO is the first publicly available dataset to combine reconstructed scenes, human activity, natural language, and contextual-placement annotations. It comprises 3,805 human-annotated episodes across 159 floors of 115 HM3D scenes, involving 198 distinct objects. Each floor includes a complete RGB-D scan with human-annotated room labels and a surface list. Each episode provides synchronized RGB-D encounter clips; 6-DoF camera, human, and object trajectories; start and destination surfaces; an action caption; and a human-written context: a single sentence describing the inhabitant's routine that implies the destination without naming it. We evaluate contextual object placement using input-masked probes and an end-to-end baseline. Results show that no single input modality is sufficient, highlighting the need to jointly reason over scene structure, human activity, and contextual knowledge.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑