arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WareFly-VLA:面向智能仓库中无人机导航与人员跟踪的视觉-语言-动作框架

WareFly-VLA: A Vision-Language-Action Framework for UAV Navigation and Human Tracking in Smart Warehouses

Thinh D. Le, Son T. Nguyen, Duong Q. Nguyen, Dung D. Le, Ngo Anh Vien, H. Nguyen-Xuan

arXiv 2610.08526首次发表:更新:

发表机构

Center for AI Research, VinUniversity; College of Engineering and Computer Science, VinUniversity; VinRobotics(VinUniversity人工智能研究中心; VinUniversity工程与计算机科学学院; VinRobotics公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对智能仓库中语言引导的无人机导航与人员跟踪问题,提出 WareFly-VLA 框架与数据集,包含 507 个飞行片段和 8,504 帧,并建立四架构基准,揭示当前 VLA 模型在空中控制中的不足。

AI 中文摘要

视觉-语言-动作(VLA)模型在机器人操作和地面移动导航中取得了显著成果,然而在智能仓库中,语言条件控制的无人机(UAV)仍 largely 未被探索,主要受限于缺乏同时提供连续低层飞行动作、细粒度自然语言目标描述以及真实工业环境的基准。本文介绍了 WareFly-VLA,一个用于仓库环境中语言引导的人员搜索、定位和跟踪的逼真无人机 VLA 框架和数据集。它包含在 NVIDIA Isaac Sim 中收集的 507 个人类远程操作飞行片段和 8,504 个高分辨率 RGB 转换,每个片段都配有人类书写的目标工人外观描述和同步的四自由度控制命令。覆盖两个空中任务:目标接近和人员跟随,在遮挡、远距离搜索、高度变化和杂乱环境下进行。在无泄漏的片段级协议下,以两种控制速率建立了四个开源 VLA 架构(SmolVLA、GR00T N1.7、pi_0 和 OpenVLA)的统一基准。结果表明,仓库中的语言条件空中控制远未解决:在严格泛化设置下性能大幅下降,连续动作建模始终优于离散动作标记化,只有前向通道可以从单帧可靠学习,当前基础模型接口从地面和人形实体到空中平台的迁移效果较差。同步的视频、语言、动作、姿态和难度注释进一步支持世界模型研究。数据集、基线和评估协议已发布,以支持智能仓库中语言基础的空中自主性。

英文摘要

Vision-Language-Action (VLA) models have achieved impressive results in robotic manipulation and ground-mobile navigation, yet language-conditioned control of unmanned aerial vehicles (UAVs) in smart warehouses remains largely unexplored, hindered by the lack of benchmarks that jointly provide continuous low-level flight actions, fine-grained natural-language target descriptions, and realistic industrial environments. This paper introduces WareFly-VLA, a photorealistic UAV VLA framework and dataset for language-guided human search, localization, and tracking in warehouse environments. It contains 507 human-teleoperated flight episodes and 8,504 high-resolution RGB transitions collected in NVIDIA Isaac Sim, each paired with a human-written appearance description of the target worker and a synchronized four-degree-of-freedom control command. Two aerial tasks are covered: target approach and person following, under occlusion, long-range search, altitude variation, and clutter. A unified benchmark of four open-source VLA architectures (SmolVLA, GR00T N1.7, pi_0 and OpenVLA) is established under a leakage-free episode-level protocol at two control rates. The results show that language-conditioned aerial control in warehouses is far from solved: performance drops substantially under strict generalization settings, continuous action modeling consistently outperforms discrete action tokenization, only the forward channel is reliably learnable from a single frame, and current foundation-model interfaces transfer poorly from ground and humanoid embodiments to aerial platforms. The synchronized video, language, action, pose, and difficulty annotations further support world-model research. The dataset, baselines, and evaluation protocol are released to support language-grounded aerial autonomy in smart warehouses.

Comments41 pages, 35 figures, 11 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑