发表机构
ShanghaiTech University; Deemos Technology(上海科技大学; 斗象科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LIVIN是基于30个有人居住住宅数字孪生的空间与具身智能基准,通过人在回路工作流构建,评估四项任务,助力具身AI在真实家庭环境发展。
AI 中文摘要
真实的家庭模拟不仅需要捕捉多样化环境,还需捕捉影响机器人运动与交互的有人居住的物体布置及空间约束。现有资源常需在规模、真实世界对应性与交互就绪性之间权衡,导致缺乏对真实住宅实际布置的忠实交互式复制品。为此,我们推出LIVIN——一个基于30个多样化有人居住住宅数字孪生的空间与具身智能基准。这些复制品保留了观测到的房间布局、家具配置及日常物品。为构建它们,我们设计了包含实例识别、建筑重建、物体生成与放置的人在回路工作流,各阶段中间结果均由人类对照源观测进行审核与修正。我们在LIVIN中评估四项任务:3D检测、3D重建、导航及 locomotion-manipulation( locomotion-manipulation 译为 locomotion-manipulation,即 locomotion 为运动,manipulation 为操作,此处保留英文原名)。评估显示,当前方法在应对真实有人居住住宅中密集的物体布置、遮挡、有限自由空间及受限交互区域时仍面临挑战。我们希望LIVIN能推动具身AI在真实住宅中的发展,从空间理解到机器人交互,最终将具身智能带入日常家庭环境。
英文摘要
Realistic household simulation must capture not only diverse environments but also the lived-in object arrangements and spatial constraints that shape robot motion and interaction. Existing resources often trade off scale, real-world correspondence, and interaction readiness, leaving a gap in faithful, interactive replicas of how real homes are actually arranged. To this end, we introduce LIVIN, a benchmark for spatial and embodied intelligence built on digital twins of 30 diverse lived-in homes. These replicas preserve observed room layouts, furniture configurations, and everyday belongings. To construct them, we design a human-in-the-loop workflow comprising instance recognition, architectural reconstruction, and object generation and placement, with intermediate results reviewed and corrected by humans against the source observations at each stage. We evaluate four tasks in LIVIN: 3D detection, 3D reconstruction, navigation, and loco-manipulation. Our evaluations show that current methods remain challenged by the dense object arrangements, occlusions, limited free space, and constrained interaction regions found in realistic lived-in homes. We hope LIVIN will help advance embodied AI in real-world homes, from spatial understanding to robotic interaction, and ultimately bring embodied intelligence into everyday home environments.