arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12860cs.RO

HumanoidVLN:面向多种人形机器人形态的基于物理的视觉语言导航模拟器与基准

HumanoidVLN: A Physics-Grounded Simulator and Benchmark for Vision-Language Navigation Across Diverse Humanoid Embodiments

发表机构VinMotion公司 · 南加州大学
查看机构详情
  • VinMotion, Inc.(VinMotion公司)
  • University of Southern California(南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

Quan-Dung Pham, Anh Dao, The-Anh Nguyen, Minh Nguyen-Dinh, Phuong Nam Dang, Tri Pham, Hung Tran, Bach Dao, Tuyen P. Le, Truong Nguyen, Quan Nguyen

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对人形机器人VLN的物理约束与形态差异问题,提出基于NVIDIA Isaac Sim的HumanoidVLN模拟器与基准,验证其与多种模型、机器人的兼容性,实验揭示了VLN模型、控制器与人形形态的交互关系。

中文摘要 AI 辅助

人形机器人的视觉语言导航(VLN)面临现有基准无法解决的挑战:双足移动带来了轮式机器人不存在的物理约束,不同平台的人形机器人形态存在差异,且自中心观测会因移动引发的相机动态而失真。我们提出HumanoidVLN,这是一个面向多种人形机器人形态的基于物理的VLN模拟器与基准。该平台基于NVIDIA Isaac Sim构建,支持可扩展的人形机器人配置,在4台机器人(Unitree G1、Unitree H1、Internal-A、Internal-B)上得到验证,这些机器人的下肢自由度(DoF)为10-12,身高为1.17m至1.80m,其分层控制栈结合了强化学习移动策略与可互换的PD或MPC路径跟踪器。新机器人与VLN模型可轻松集成,我们验证了其与NaVILA、DualVLN、StreamVLN和JanusVLN的兼容性。环境来自艺术家设计的场景与3D Gaussian Splatting重建,筛选出可导航区域超过100平方米的场景。指令由双生成-审核+释义多智能体流水线生成,并经人工验证,共得到933个碰撞感知参考episode,每个episode对应1条细粒度指令和3条粗粒度风格变体(正式、自然、随意)。在4种模型和4种形态下,JanusVLN实现最高平均成功率43.55%,nDTW为48.38。在DualVLN与Unitree G1的20集仿真到真实迁移试点中,导航误差相关性极强(r=0.935),平均绝对误差为0.68m,平均轨迹相似度为0.782(±0.188)nDTW。这些结果凸显了VLN模型、控制器与人形机器人形态在物理执行下的相互作用。代码、基准与数据将在录用后于该httpsURL发布。

英文摘要

Vision-Language Navigation (VLN) for humanoid robots poses challenges existing benchmarks fail to address: bipedal locomotion imposes physical constraints absent from wheeled agents, humanoid morphologies vary across platforms, and egocentric observations are distorted by locomotion-induced camera dynamics. We present HumanoidVLN, a physics-grounded simulator and benchmark for VLN across diverse humanoid embodiments. Built on NVIDIA Isaac Sim, our platform supports an extensible set of humanoid configurations, demonstrated on four robots (Unitree G1, Unitree H1, Internal-A, Internal-B) spanning 10-12 lower-body DoF and heights from 1.17m to 1.80m, via a hierarchical control stack combining a reinforcement learning locomotion policy with interchangeable PD or MPC path trackers. New robots and VLN models integrate with minimal effort; we demonstrate compatibility with NaVILA, DualVLN, StreamVLN, and JanusVLN. Environments are drawn from artist-designed scenes and 3D Gaussian Splatting reconstructions, filtered for navigable areas exceeding 100 square meters. Instructions are generated by a dual generator-reviewer plus paraphraser multi-agent pipeline with human-in-the-loop verification, yielding 933 collision-aware reference episodes, each paired with one fine-grained instruction and three coarse-grained stylistic variants (formal, natural, casual). Across four models and four embodiments, JanusVLN achieves the highest mean success rate of 43.55% and nDTW of 48.38. In a 20-episode sim-to-real pilot with DualVLN and the Unitree G1, navigation errors correlate strongly (r=0.935), with a mean absolute difference of 0.68m and mean trajectory similarity of 0.782 (+/-0.188) nDTW. These results highlight the interaction between VLN models, controllers, and humanoid embodiments under physical execution. Code, benchmark, and data will be released upon acceptance at https://humanoid-vln.github.io/.

相关深度报道

↑