CrossTracer:基于VLA模型推理与轨迹残差自适应的跨 embodiments 导航
CrossTracer: Cross-Embodiment Navigation via VLA Model Reasoning and Trace Residuals Adapting
浏览论文内容
中文总结 AI 辅助
本研究提出跨 embodiment 导航框架 CrossTracer,通过 VL-Tracer 和 CE-Adapter 优化轨迹,在 NaviTrace 基准得分 45.68,优于 Gemini-2.5-Pro,实际部署于轮式、腿式机器人效果良好。
中文摘要 AI 辅助
视觉-语言-动作(Vision-Language-Action,VLA)模型为机器人导航提供了强大的语义先验,但往往忽略特定 embodiment 的移动约束。对某一机器人语义合理的路径,对另一机器人可能在物理上不可行。我们提出 CrossTracer,这是一种通过自适应轨迹残差实现跨 embodiments 导航的分层框架。CrossTracer 将导航计划表示为归一化图像平面路标点,形成语义推理与物理接地之间的统一像素空间接口。首先,视觉-语言轨迹提议器(Vision-Language Trace Proposer,VL-Tracer)对预训练 VLA 模型进行适配,使其能从自我中心观测和灵活目标规格预测初始导航轨迹。其次,CE-Adapter 通过从视觉可通行性线索、机器人身份和初始轨迹预测 embodiment 条件残差修正来优化该轨迹。为在无需昂贵手动标注的情况下训练优化模块,跨 embodiment RRT*(Cross-Embodiment RRT*,CE-RRT*)将全景分割转换为机器人条件可通行性代价图,并生成代价最小化的像素空间轨迹。我们在 NaviTrace 基准上评估 CrossTracer,该基准测试模型能否从自我中心观测、语言指令和机器人 embodiment 类型生成与 embodiment 一致的导航轨迹。CrossTracer 总得分达 45.68,比评估的最强通用基线 Gemini-2.5-Pro 高出 10.01 分,对应相对提升 28.1%。在轮式和腿式机器人上的实际部署进一步显示导航成功率和执行效率得到提升。
英文摘要
Vision-language-action (VLA) models provide strong semantic priors for robot navigation, but they often ignore embodiment-specific mobility constraints. A path that is semantically plausible for one robot may be physically infeasible for another. We propose CrossTracer, a hierarchical framework for cross-embodiment navigation through adaptive trace residuals. CrossTracer represents navigation plans as normalized image-plane waypoints, forming a unified pixel-space interface between semantic reasoning and physical grounding. First, Vision-Language Trace Proposer (VL-Tracer) adapts a pretrained VLA model to predict an initial navigation trace from egocentric observations and flexible goal specifications. Second, CE-Adapter refines this trace by predicting embodiment-conditioned residual corrections from visual traversability cues, robot identity, and the initial trace. To train the refinement module without costly manual annotation, Cross-Embodiment RRT* (CE-RRT*) converts panoptic segmentation into robot-conditioned traversability cost maps and generates cost-minimizing pixel-space traces. We evaluate CrossTracer on the NaviTrace benchmark, which tests whether a model can generate embodiment-consistent navigation traces from egocentric observations, language instructions, and robot embodiment types. CrossTracer achieves a total score of 45.68, outperforming the strongest evaluated general-purpose baseline, Gemini-2.5-Pro, by 10.01 points, corresponding to a 28.1% relative improvement. Real-world deployment on wheeled and legged robots further shows improved navigation success and execution efficiency.
发表机构
- Peng Cheng Laboratory(鹏城实验室)
- Southern University of Science and Technology(南方科技大学)
- Innovation Investment Research Institute(创新投资研究院)
- Soochow University(苏州大学)
机构由 AI 辅助整理,请以论文原文为准。