发表机构
Rekise Marine(瑞基斯海洋)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对开放词汇3D地图无法处理场景元素运动查询的问题,提出视觉-语言-运动地图(VLMM),其元素有融合运动属性及不确定性。通过实验验证模式字段不可替代,在真实数据集上该方法提升精度、减少错误标志且对噪声姿态鲁棒。
AI 中文摘要
开放词汇的3D地图能让机器人回答关于物体位置的语言查询,但假定世界是静态的,无法回答关于场景元素如何移动的问题。我们引入了视觉-语言-运动地图(VLMM),这是一种开放词汇、可通过自然语言查询的3D地图,其中每个元素都带有融合的运动属性:结合了几何观测到的跨帧运动的VLM/LLM语义可移动性先验,以及每个元素的不确定性。查询简化为属性过滤器,可区分已观察到移动的、可能移动但未移动的以及静止不动的物体。在具有精确地面真值的受控模拟器基准测试(AI2-THOR,三种场景类型)中,通过消融实验表明模式字段不可替代:仅语义的基线即使有强大特征也无法回答运动查询,且两个运动字段不能相互替代。在真实动态RGB-D数据集(TUM和Bonn,六个序列)上,我们展示了不确定性通道——我们与先前融合运动工作的关键区别——持续提高了移动与静态的平均精度,减少了错误运动标志,并且对估计的(有噪声的)姿态具有鲁棒性。原始置信度未校准,但事后等渗校准达到了0.10的预期校准误差。VLMM是一种表示贡献:最接近的先前地图各自至少缺少我们组合提供的四个属性中的一个——开放词汇、语言可查询、融合先验和观测运动以及每个元素的不确定性。
英文摘要
Open-vocabulary 3D maps let robots answer language queries about what and where, but they assume a static world and cannot answer queries about how scene elements behave. We introduce Vision-Language-Motion Maps (VLMM), an open-vocabulary, language-queryable 3D map - queried through a rule-based intent router over open-vocabulary object nouns, not a general natural-language interface - in which each element carries a fused motion attribute: a VLM/LLM semantic movability prior combined with geometrically observed cross-frame motion, together with a per-element uncertainty. Queries reduce to attribute filters that distinguish what has been seen to move, what could move but has not, and what stays still. On a controlled simulator benchmark with exact ground truth (AI2-THOR, three scene types) we show through ablation that the schema fields are non-substitutable: a semantic-only baseline fails motion queries even with strong features, and neither motion field substitutes for the other (the prior cannot answer "what is moving," observed motion cannot answer "what could move"). On real dynamic RGB-D (TUM and Bonn, six sequences) we show the uncertainty channel - our key difference from prior fused-motion work - consistently improves moving-vs-static average precision and reduces false motion flags, and that it is robust to estimated (noisy) poses. The raw confidence is not calibrated, but post-hoc isotonic calibration reaches an expected calibration error of 0.10. VLMM is a representation contribution: the closest prior maps each lack at least one of the four properties - open-vocabulary, language-queryable, fused prior-and-observed motion, and per-element uncertainty - that our combination provides.
Comments8 pages, 5 figures, 3 tables. v2: corrected Eq. (7) (residual covariance;implementation unaffected), added the persistent-map update rule, the query-parser grammar, and an aggregation-quantile ablation