arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

助手放置基准Aria:面向以自我为中心的放置辅助任务的基准

Assistant Placement Aria: A Benchmark for Egocentric Placement Assistance

Amir Belder, Gonçalo Dias Pais, Refael Vivanti, Omri Carmi, Daniel DeTone, Oren Shrout, Ido Gattegno, Ayellet Tal

arXiv 2608.00652首次发表:更新:

发表机构

Reality Labs, Meta inc.; Technion, institute of technology; Sensei(Meta公司现实实验室; 以色列理工学院; Sensei)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出首个虚拟放置基准Assistant Placement Aria,涵盖三类以人为中心的放置任务,在基准上评估了多个基础模型,推动虚拟放置领域研究。

AI 中文摘要

机器人领域的人类辅助涉及导航、物体操作、放置等多个任务,其中关键挑战是选择符合人类意图或偏好的目标位置。本文聚焦虚拟放置(Virtual Placement, VP)场景下的该挑战,虚拟放置是指结合场景上下文与以人为中心的约束,识别所有合理目标位置的任务,这与传统放置任务通常聚焦单一预定义目标位置的模式不同。虚拟放置问题较为复杂,需对场景的几何、语义及合理性进行全局与局部推理。为填补该空白,本文提出Assistant Placement Aria,这是首个探索虚拟放置多方面内容的基准,涵盖全局、局部及以人为中心的约束。该基准包含合成与真实室内场景,标注了三项任务:(i)2D面板放置、(ii)就座建议、(iii)电视放置;每个场景包含2D图像、3D点云及场景内物体的文本描述。本文发布该基准,旨在推动对这一依赖相关数据的未充分探索且具挑战性领域的进一步研究,同时在基准上评估了多个用于物体检测与分割的基础模型。

英文摘要

Human assistance in robotics spans around several tasks such as navigation, object manipulation, and placement, where a key challenge is selecting target destinations that align with human intentions or preferences. We focus on this challenge in the context of Virtual Placement (VP), the task of identifying all plausible target locations given scene context and human-centric constraints. This differs from traditional placement tasks that typically focus on a single, predefined target location. The VP problem is complex, as it requires both global and local reasoning about the scene's geometry, semantics, and plausibility. To address this gap, we introduce {\bf Assistant Placement Aria}, the first benchmark to explore diverse aspects of VP, including global, local, and human-centric constraints. It contains both synthetic and real indoor scenes annotated for three tasks: (i)~2D Panel Placement, (ii)~Sitting Suggestion, and (iii)~TV Placement. Each scene includes 2D images, a 3D point cloud, and a textual description of the objects within the scene. By contributing this benchmark, we aim to encourage further research in this underexplored and challenging field that is critically dependent on relevant data. We also evaluate several foundation models for object detection and segmentation on our benchmark.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑