CognitiveReality:基于LLM智能体的机器人无关语义高斯建图,用于沉浸式协作VR遥操作
CognitiveReality: Robot-Agnostic Semantic Gaussian Mapping with an LLM Agent for Immersive Collaborative VR Teleoperation
- Skolkovo Institute of Science and Technology(斯科尔科沃科学技术学院)
- NLP Research Center(自然语言处理研究中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
CognitiveReality通过LLM智能体将机器人RGB-D流转化为语义高斯地图,实现VR遥操作中的工具使用与语音交互,在四足机器人上达到高成功率与建图质量。
AI中文摘要:
逼真的三维视图能告诉遥操作员机器人的位置,但不能告诉场景中包含什么、每个物体被观测得如何,或如何将指向和语音转化为机器人动作。CognitiveReality将机器人的RGB-D流转化为实时、语义索引的高斯-截断符号距离函数(TSDF)地图,由虚拟现实中的操作员和使用工具的语言智能体共享。一个建图二进制文件仅通过配置即可服务于任何平台:它从机器人SLAM、关节运动学、动作捕捉或内联视觉跟踪器接收位姿,通过阴影跟踪器和关键帧锚定的PnP(透视n点)算法弥补定位中断,并以2 Hz的频率维护具有逐物体质量的开词汇实例身份。语音和控制器射线通过验证的具类型工具和操作员确认的机器人动作,与持久场景物体建立关联。在受控智能体评估中,部署的本地Qwen3-VL-8B路由器达到81.24%的工具精确匹配,而合并感知重放正确重定向了101个被吸收的物体标识符。在机器人数据上,CognitiveReality比高斯加SDF基线高出2-8 dB;在5-40秒的SLAM中断期间,位姿误差保持在1-8厘米内。在两台四足机器人上实时部署,智能体执行了30个导航请求中的26个和20个重新观测请求中的20个,将物体质量提高了2-5 dB。
英文摘要:
A photorealistic 3D view tells a teleoperator where a robot is, but not what the scene contains, how well each object has been observed, or how to turn pointing and speech into robot action. CognitiveReality turns a robot's RGB-D stream into a live, semantically indexed Gaussian-TSDF map shared by an operator in virtual reality and a tool-using language agent. One mapper binary serves any platform through configuration alone: it ingests poses from robot SLAM, joint kinematics, motion capture or an inline visual tracker, bridges localization outages with a shadow tracker and keyframe-anchored PnP, and maintains open-vocabulary instance identities with per-object quality at 2 Hz. Speech and controller rays are grounded against persistent scene objects through validated typed tools and operator-confirmed robot actions. In the controlled agent evaluation, the deployed local Qwen3-VL-8B router reaches 81.24\% tool exact match, while merge-aware replay correctly redirects 101 absorbed object identifiers. On robot data CognitiveReality exceeds a Gaussian-plus-SDF baseline by 2-8 dB; pose error through 5-40 s SLAM outages stays within 1-8 cm. Deployed live on two quadrupeds, the agent executed 26 of 30 navigation requests and 20 of 20 re-observation requests, raising object quality by 2-5 dB.