发表机构
Dublin City University; Insight Research Ireland Centre for Data Analytics; South East Technological University(都柏林城市大学; 爱尔兰洞察数据分析研究中心; 东南理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一个整合LLM、VLM和数字孪生的端到端框架,利用手机摄像头和SLAM3R生成3D点云,经SpatialLM和本地LLM处理,为视障用户提供精确的空间认知导航支持,实验验证了其准确性。
AI 中文摘要
由大型语言模型(LLMs)和视觉-语言模型(VLMs)驱动的多模态人工智能,通过实现视觉和文本数据的同步处理,正在变革辅助技术领域。这一进展对全球超过4300万视障和神经多样性个体具有重大意义,这些个体因空间意识有限和环境线索不足,在室内外环境导航中面临持续挑战。现有的导航辅助工具往往缺乏全面的3D场景理解,依赖受限的基于路线的策略,这阻碍了用户的自主性。在本文中,我们提出了一种新颖的端到端框架,该框架整合了LLMs、VLMs和数字孪生技术,为视障和神经多样性用户提供空间认知导航支持。我们的系统通过标准手机摄像头捕获视频输入,并采用SLAM3R从单目RGB序列中实时生成密集3D点云。我们定制的后处理算法确保在无需预定义参考点的情况下,跨多个视角实现精确的点云对齐。这增强了SpatialLM生成结构化3D表示的能力,包括建筑元素和定向对象边界框。随后,丰富后的空间数据由本地部署的LLM处理,该LLM解释3D上下文以生成详细的场景描述以及用户与周围物体之间的精确距离测量。我们在多种视频场景中评估了我们的方法,这些场景涵盖不同视角、环形行走视图,并在多种环境中捕获。评估结果表明,在3D场景解释和对象定位方面具有一致的准确性,凸显了我们的系统作为一种变革性辅助导航解决方案的潜力,该方案将先进的视觉感知与空间推理相结合。
英文摘要
Multimodal AI, powered by Large Language Models (LLMs) and Vision-Language Models (VLMs), is transforming assistive technologies by enabling simultaneous processing of visual and textual data. This advancement holds significant promise for over 43 million visually impaired and neuro-divergent individuals worldwide who face persistent challenges in navigating indoor and outdoor environments due to limited spatial awareness and insufficient environmental cues. Existing navigation aids often lack comprehensive 3D scene understanding, relying on constrained route-based strategies that hinder user autonomy. In this paper, we introduce a novel end-to-end framework that integrates LLMs, VLMs and digital twin technologies to deliver a spatially cognitive navigation support for visually impaired and neuro-divergent users. Our system captures video input via standard mobile phone cameras, and employs SLAM3R to generate dense 3D point clouds from monocular RGB sequences in real-time. Our custom post-processing algorithm ensures accurate point cloud alignment across multiple viewpoints without requiring predefined reference points. This enhances the capabilities of SpatialLM to produce structured 3D representations, including architectural elements and oriented object bounding boxes. The enriched spatial data is then processed by a locally deployed LLM, which interprets 3D contexts to generate detailed scene descriptions and precise distance measurements between users and surrounding objects. We evaluated our approach across diverse video scenarios featuring various perspectives, looped walking views and captured in multiple environments. The evaluation results demonstrate consistent accuracy in 3D scene interpretation and object localisation, underscoring the potential of our system as a transformative assistive navigation solution that combines advanced visual perception with spatial reasoning