发表机构
The Hong Kong University of Science and Technology (Guangzhou); Southern University of Science and Technology; Capstone Technology Co., Ltd.(香港科技大学(广州); 南方科技大学; 凯普斯通科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ExcavaTwin提出一种无需训练的纯视觉几何引导语义高程映射框架,通过冻结视觉模型重建几何与语义,并利用几何约束融合多视角信息,实现自主挖掘中可靠的地形感知,实验验证其有效性与实时性。
AI 中文摘要
自主挖掘需要一种空间表示,该表示能同时捕捉地形几何和与任务相关的语义信息。现有的挖掘映射主要侧重于高程,而通用的语义模型在非结构化户外场景中仍不稳定。我们提出了ExcavaTwin,一种无需挖掘特定训练的纯视觉几何引导语义高程映射框架。给定多视角RGB图像,该框架:1)使用冻结的视觉模型重建场景几何和语义观测;2)推导地形和非地形的几何支撑;3)执行几何约束的多视角语义融合,以抑制不合理的预测并恢复不完整的观测;4)将融合状态投影到面向任务的语义高程地图中。在公共数据集和真实挖掘场景上的实验证明了可靠的几何和语义感知。在真实挖掘中,系统实现了约1.4秒的平均更新间隔,在动态修改区域的平均高程误差为12.74厘米。较大的误差主要发生在快速地形变化和机器运动引起的瞬时视觉干扰期间。
英文摘要
Autonomous excavation requires a spatial representation that jointly captures terrain geometry and task-relevant semantics. Existing excavation mapping is largely elevation-centric, while generic semantic models remain unstable in unstructured outdoor scenes. We present ExcavaTwin, a pure-vision geometry-guided semantic elevation mapping framework without excavation-specific training. Given multi-view RGB images, the framework: 1) reconstructs scene geometry and semantic observations using frozen vision models; 2) derives terrain and non-terrain geometric support; 3) performs geometry-constrained multi-view semantic fusion to suppress implausible predictions and recover incomplete observations; and 4) projects the fused state into a task-oriented semantic elevation map. Experiments on public datasets and real excavation scenes demonstrate reliable geometric and semantic perception. In real excavation, the system achieved an average update interval of approximately 1.4 s and a mean elevation error of 12.74cm in dynamically modified regions. Larger errors mainly occur during rapid terrain changes and transient visual disturbances caused by machine motion.
Comments8 pages,5 figures