发表机构
Intelligent Space Robotics Laboratory, Skolkovo Institute of Science and Technology; NLP Research Center(智能空间机器人实验室,斯科尔科沃科学技术研究院; 自然语言处理研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出用于开放词汇VR场景探索的集成架构,结合ORB-SLAM3姿态估计与在线高斯重建处理现实数据,通过VR管道探索,语义模块转录语音等,在数据集和TUM-RGBD上提升图像质量,帧率可比或更优,实现88%的视觉语言模型对象识别率。
AI 中文摘要
我们提出了一种新颖的集成架构,用于鲁棒在线3D高斯点云、实时VR探索和语音驱动的视觉语言模型交互。与假设深度或外部姿态干净的方法不同,我们的系统将基于ORB-SLAM3的姿态估计与在线高斯重建相结合,用于嘈杂现实世界数据。VR管道实现了对增量重建的沉浸式探索;语义模块转录语音命令,生成场景描述并记录兴趣点。与现有在线高斯点云方法相比,我们在数据集和TUM-RGBD上提高了图像质量,帧率相当或更高,视觉语言模型对象识别率达到88%。
英文摘要
We present a novel integrated architecture for robust online 3D Gaussian splatting, real-time VR exploration, and speech-driven Vision-Language-Model interaction. Unlike methods assuming clean depth or external poses, our system combines ORB-SLAM3-based pose estimation with online Gaussian reconstruction for noisy real-world data. A VR pipeline enables immersive exploration of incremental reconstructions; a semantic module transcribes voice commands, generates scene descriptions, and records points of interest. Against state-of-the-art online Gaussian splatting methods, we improve image quality on our dataset (+14.5% PSNR, +8.6% SSIM, -14.3% LPIPS) and TUM-RGBD (+11.7% PSNR, +7.8% SSIM, -21.6% LPIPS), with comparable or superior frame rates via quality-speed configurations. We achieve an 88% VLM object-recognition rate.