arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2510.22672cs.CVcs.CLcs.RO

Look and Tell:一个用于跨自我中心视角与外中心视角多模态接地的数据集

Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views

  • KTH Royal Institute of Technology(皇家理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Anna Deichler, Jonas Beskow

更新

AI总结:

本研究推出多模态数据集Look and Tell,通过智能眼镜与固定摄像头采集25名参与者厨房指称交际的多模态同步数据,结合3D重建构建跨视角空间表征评估基准,助力具身智能体的情境对话能力研发。

AI中文摘要:

我们推出Look and Tell,这是一个用于研究自我中心与外中心视角下指称交际的多模态数据集。我们使用Meta Project Aria智能眼镜与固定摄像头,记录了25名参与者指导搭档识别厨房中食材时的同步注视、语音与视频数据。结合3D场景重建,该设置为评估不同空间表征(2D与3D;自我中心与外中心)如何影响多模态接地提供了基准。该数据集包含3.67小时的录制内容,涵盖2707条标注丰富的指称表达,旨在推动具身智能体的发展,使其能够理解并参与情境对话。

英文摘要:

We introduce Look and Tell, a multimodal dataset for studying referential communication across egocentric and exocentric perspectives. Using Meta Project Aria smart glasses and stationary cameras, we recorded synchronized gaze, speech, and video as 25 participants instructed a partner to identify ingredients in a kitchen. Combined with 3D scene reconstructions, this setup provides a benchmark for evaluating how different spatial representations (2D vs. 3D; ego vs. exo) affect multimodal grounding. The dataset contains 3.67 hours of recordings, including 2,707 richly annotated referential expressions, and is designed to advance the development of embodied agents that can understand and engage in situated dialogue.

补充信息

↑