发表机构
Leipzig University; Systems Research Institute Polish Academy of Sciences; Wrocław University of Economics(莱比锡大学; 波兰科学院系统研究所; 弗罗茨瓦夫经济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出用语义辐射场(SRF)构建兼具几何真实性与语义可查询性的真实场景模拟器,用于训练评估具身智能体的空间推理能力,还将其应用于果园苹果采摘任务。
AI 中文摘要
训练和评估具身智能体的空间推理能力,需要同时具备几何真实性和语义可查询性的多样化环境。合成模拟器提供真实语义但缺乏真实感;基于真实环境重建的模拟器虽具备真实外观,但默认缺乏真实语义。我们提出使用语义辐射场(Semantic Radiance Fields, SRF)作为空间推理智能体的模拟器。SRF是一种统一上述需求的表示形式,它将预训练视觉模型的2D语义分割提升为3D辐射场,该场共同编码几何、外观和每类语义身份。所得场从真实场景的已位姿RGB捕获数据重建,支持新视角合成、语义和自由空间查询,且全部在单一基础表示中完成。这使得能够高效生成多样化真实环境,用于训练和评估空间推理模型。作为示例应用,我们概述了一个SRF驱动的果园苹果采摘任务模拟器,其中辐射场为物理引擎提供相机渲染、语义真实值和占用查询。
英文摘要
Training and evaluating spatial reasoning in embodied agents requires diverse environments that are both geometrically faithful and semantically queryable. Synthetic simulators offer ground truth semantics but sacrifice realism; simulators based on reconstructions of real-world environments have realistic appearance but lack ground truth semantics by default. We propose using Semantic Radiance Fields (SRF) as simulators for spatial reasoning agents. SRFs are a representation that unifies these requirements by lifting 2D semantic segmentations from pretrained vision models into a 3D radiance field that jointly encodes geometry, appearance, and per-class semantic identity. The resulting fields are reconstructed from posed RGB captures of real scenes and support novel-view synthesis, semantic and free-space queries within a single grounded representation. This enables the efficient generation of diverse real-world environments to train and evaluate spatial reasoning models. As an example application, we outline an SRF-driven simulator for an orchard apple-reaching task, in which the radiance field supplies camera rendering, semantic ground truth, and occupancy queries to a physics engine.
CommentsAccepted at the IJCAI 2026 Workshop on Spatio-Temporal Reasoning and Learning (STRL), oral presentation