AI 中文总结
LightNav-0 是激发 VLM 空间智能的通用具身导航模型,单模型支持多导航任务,在仿真基准和真实场景中均实现最优性能,具备强泛化能力。
AI 中文摘要
具身导航要求智能体在不同任务、环境和机器人 embodiment(形态)下,将异构目标与视觉观测转化为动作。现代视觉语言模型(VLM)已编码用于视觉 grounding( grounding 指视觉定位)、空间推理和指向的空间先验,但这些能力很少被直接用于机器人控制。现有导航系统依赖特定任务或特定 embodiment 的组件,割裂了感知、推理和动作,泛化能力有限。本文提出 LightNav-0,一种紧凑的通用具身导航模型,它激发预训练 VLM 的空间智能并使其与导航对齐,无需特定任务的预测头。LightNav-0 通过统一 token 接口表示多样导航任务:双通道指向表达任务、场景和 embodiment 无关的空间意图,残差向量量化动作 tokenizer 将该意图映射到精确的特定 embodiment 轨迹。结合时间感知视觉历史压缩、ER 中间训练、监督微调与强化学习,该框架支持指令跟随、开放词汇对象导航和视觉跟踪,且所有任务在单一模型中完成。导航训练语料库涵盖 2000+ 场景和 4000+ 小时具身导航数据。用于初始化 LightNav-0 的具身推理检查点 LightNav-ER 在 8 个具身推理基准中取得最高完整集平均性能,而 LightNav-0 在全部 10 个公开导航仿真设置中达到最先进的单目成功率。真实世界评估进一步证明其在机器人 embodiment、多样场景以及静态和动态目标上的零样本泛化能力。这些结果确立了紧凑 VLM 作为通用具身导航的统一可迁移主干。
英文摘要
Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across tasks, environments, and robot embodiments. Modern vision-language models (VLMs) already encode spatial priors for visual grounding, spatial reasoning, and pointing, but these capabilities are rarely elicited directly for robot control. Existing navigation systems instead rely on task- or embodiment-specific components, fragmenting perception, reasoning, and action while offering limited generalization. Here we present LightNav-0, a compact generalist embodied navigation model that elicits the spatial intelligence of a pretrained VLM and aligns it with navigation, without task-specific prediction heads. LightNav-0 represents diverse navigation tasks through a unified token interface: dual-channel pointing expresses task-, scene-, and embodiment-agnostic spatial intent, while a residual vector-quantized action tokenizer maps this intent to precise, embodiment-specific trajectories. Together with temporally aware visual history compression, ER mid-training, supervised fine-tuning, and reinforcement learning, this formulation supports instruction following, open-vocabulary object navigation, and visual tracking within a single model. The navigation training corpus spans 2K+ scenes and 4K+ hours of embodied navigation data. LightNav-ER, the embodied-reasoning checkpoint used to initialize LightNav-0, attains the highest complete-set average across 8 embodied-reasoning benchmarks, while LightNav-0 achieves state-of-the-art monocular success rates across all 10 public navigation simulation settings. Real-world evaluations further demonstrate zero-shot generalization across robot embodiments, diverse scenes, and static and dynamic targets. These results establish compact VLMs as a unified and transferable backbone for generalist embodied navigation.
CommentsTechnical report