发表机构
College of Integrated Circuits, Hunan University; National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University; School of Advanced Technology, Xi’an Jiaotong-Liverpool University; Suzhou VSDeep Intelligent Technology Co., Ltd.(湖南大学集成电路学院; 湖南大学机器人视觉感知与控制技术国家工程研究中心; 西交利物浦大学智能工程学院; 苏州维深智能科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出UniTrackPLA统一全景-语言-动作模型,通过全景感知编码和世界-动作一致性实现指令导航与动态人员跟踪,显著提升成功率并引入OmniTrackNav-Bench基准。
AI 中文摘要
通用型具身机器人应同时支持向语言指定目的地导航以及在任意初始目标方位角下进行动态人员跟踪。然而,现有方法通常依赖前视观测,并通过独立的策略来处理这些任务,这限制了全方位感知和统一的闭环控制。我们提出了UniTrackPLA,一个用于指令引导导航和动态人员跟踪的统一全景-语言-动作模型。其全景感知编码(PAE)保留了从每个全景图投影出的透视视图的时间和方位角结构,使预训练的透视视觉编码器能够处理全方位观测。一个共享的视觉-语言骨干网络将指令置于全景上下文中,并为两个任务预测连续的以机器人为中心的航点块。世界-动作一致性(WAC)进一步预测动作条件下的未来视觉状态,并在线验证航点前缀,允许在出现不一致时重用可靠的动作并触发重新规划。我们还引入了OmniTrackNav-Bench,包含5,000条模拟跟踪轨迹、10,000条模拟VLN路线和96条经过验证的真实世界路线,提供了919,978个航点监督实例。UniTrackPLA将整体跟踪成功率从23.50%提高到35.00%,并将Omni-VLN SR/SPL从13.00%/12.77%提高到19.75%/19.29%。纳入76条真实世界路线进一步将保留的EP@0.2m从42.92%提高到92.08%。在Go2-W机器人上的闭环实验证明了在室内和室外环境中统一的全景跟踪和导航能力。项目页面位于此https URL。
英文摘要
General-purpose embodied robots should support both navigation toward language-specified destinations and dynamic person tracking under arbitrary initial target azimuths. However, existing methods typically rely on forward-facing observations and address these tasks with separate policies, limiting omnidirectional perception and unified closed-loop control. We present UniTrackPLA, a unified panorama-language-action model for instruction-guided navigation and dynamic person tracking. Its Panoramic-Aware Encoding (PAE) preserves the temporal and azimuthal structure of perspective views projected from each panorama, enabling perspective-pretrained visual encoders to process omnidirectional observations. A shared vision-language backbone grounds instructions in the panoramic context and predicts continuous robot-centric waypoint chunks for both tasks. World-Action Consistency (WAC) further predicts action-conditioned future visual states and verifies waypoint prefixes online, allowing reliable actions to be reused while triggering replanning upon inconsistency. We also introduce OmniTrackNav-Bench, comprising 5,000 simulated tracking trajectories, 10,000 simulated VLN routes, and 96 verified real-world routes, providing 919,978 waypoint-supervision instances. UniTrackPLA improves overall tracking SR from 23.50% to 35.00% and Omni-VLN SR/SPL from 13.00%/12.77% to 19.75%/19.29%. Incorporating 76 real-world routes further improves held-out EP@0.2m from 42.92% to 92.08%. Closed-loop experiments on a Go2-W robot demonstrate unified panoramic tracking and navigation across indoor and outdoor environments. The project page is at https://tw5775.github.io/UniTrackPLA.
CommentsThe project page is at https://tw5775.github.io/UniTrackPLA