场景语言模型用于开放词汇场景映射
A Scene Language Model for Open-Vocabulary Scene Mapping
浏览论文内容
中文总结 AI 辅助
SceneLM是一种场景语言模型,直接维护文本场景地图,通过添加、编辑和删除对象更新地图,在检索和定位基准上性能与完整系统相当,且表示紧凑6-12倍,可在线运行于边缘设备。
中文摘要 AI 辅助
开放词汇的3D场景映射旨在构建环境中对象的持久表示。现有系统通常依赖工程化的映射管线来关联观测、跨视图合并信息,并随时间维护一致的场景表示。许多系统还存储富含特征的对象表示,如嵌入或图像裁剪,这增加了持久内存的大小和复杂性。我们引入了SceneLM,一种场景语言模型,直接维护文本场景地图。整个场景被表示为对象的结构化文本列表,作为模型唯一的持久内存。对于每个输入图像,模型读取当前场景状态并通过添加、编辑和删除对象来更新地图。为了学习这种行为,我们引入了迭代场景地图维护的监督任务,以及一个自动标注管线,该管线从图像生成训练数据而无需人工标签。我们在基于语言锚定的检索基准和定位基准上评估了SceneLM。在两个基准上,该模型生成的场景地图在性能上与由专用感知和几何模块构建的完整映射系统相当,同时场景表示紧凑6-12倍。我们进一步通过四足机器人上的实验表明,SceneLM可以在边缘设备上在线运行。这些结果表明,持久开放词汇的3D场景地图可以由单个视觉语言模型仅使用轻量级文本表示直接维护。训练和推理代码可在该https URL上获取。
英文摘要
Open-vocabulary 3D scene mapping aims to build a persistent representation of the objects in an environment. Existing systems typically rely on engineered mapping pipelines to associate observations, merge information across views, and maintain a consistent scene representation over time. Many additionally store feature-rich object representations, such as embeddings or image crops, increasing the size and complexity of the persistent memory. We introduce SceneLM, a Scene-Language Model that directly maintains a textual scene map. The full scene is represented as a structured text list of objects, which serves as the model's only persistent memory. For each input image, the model reads the current scene state and updates the map by adding, editing, and removing objects. To learn this behavior, we introduce supervision tasks for iterative scene map maintenance together with an automatic annotation pipeline that generates training data from images without human labels. We evaluate SceneLM on both a language-grounded retrieval benchmark and a localization benchmark. Across both benchmarks, the model produces a scene map that achieves competitive performance with complete mapping systems built from dedicated perception and geometric modules while producing a scene representation that is 6-12x more compact. We further show that SceneLM can be run online on an edge device through experiments on a quadruped. These results show that a persistent open-vocabulary 3D scene map can be maintained directly by a single vision-language model using only a lightweight text representation. Training and inference code is available on https://goldengait.github.io/scenelm/.
发表机构
- Chalmers(查尔姆斯理工大学)
- Zenseact
- Stanford(斯坦福大学)
- UC Berkeley(加州大学伯克利分校)
- NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。