AeroManip-VLA:基于强化学习生成演示的可扩展空中操纵视觉-语言-动作学习
AeroManip-VLA: Scalable Vision-Language-Action Learning for Aerial Manipulation with RL-Generated Demonstrations
浏览论文内容
中文总结 AI 辅助
提出AeroManip-VLA,一个基于GPU加速仿真和强化学习自动生成演示的空中操纵VLA基准,支持数据生成、轨迹分析与策略评估。
中文摘要 AI 辅助
空中操纵器将机器人操纵扩展到地面机器人难以进入的三维工作空间,为通用操纵创造了新的机会。然而,将视觉-语言-动作(VLA)模型扩展到空中机器人带来了独特的挑战,这源于操纵与飞行的紧密耦合、持续变化的观测以及安全关键的物理交互。这些挑战需要多样化的训练数据和系统性的策略评估,但在物理空中平台上收集演示和评估策略成本高昂、难以扩展,且难以在受控条件下重复。我们提出了AeroManip-VLA,一个用于空中VLA数据生成和策略评估的可扩展基准。AeroManip-VLA提供了一个GPU加速的仿真框架,在超大规模并行环境中具备低层级、负载感知的飞行和操纵控制。基于该框架,我们将可复用的强化学习策略与专家任务规则相结合,在多样化的物体、环境和随机初始条件下自动生成演示,无需人工遥操作。生成的数据包括抓取和放置等基本技能,以及需要导航和操纵的长时程任务。我们进一步引入了自动化事件标注和轨迹分类来过滤演示。这些机制能够对任务进展、行为结果和安全相关失败进行细粒度分析。最后,我们在不同任务设置下评估了一系列模仿学习和VLA基线,揭示了它们的性能特征和失败模式。综上,AeroManip-VLA在真实世界部署之前,能够在仿真中实现可扩展的空中操纵数据生成、结构化轨迹分析和系统性VLA评估。
英文摘要
Aerial manipulators extend robotic manipulation into 3D workspaces that are difficult for ground-based robots to access, creating new opportunities for general-purpose manipulation. However, extending Vision-Language-Action (VLA) models to aerial robots introduces distinct challenges due to the tight coupling between manipulation and flight, continuously changing observations, and safety-critical physical interactions. These challenges demand diverse training data and systematic policy evaluation, yet collecting demonstrations and evaluating policies directly on physical aerial platforms are costly, difficult to scale, and hard to repeat under controlled conditions. We present AeroManip-VLA, a scalable benchmark for aerial VLA data generation and policy evaluation. AeroManip-VLA provides a GPU-accelerated simulation framework with low-level payload-aware flight and manipulation control in massively parallel environments. Building on this framework, we combine reusable reinforcement learning policies with expert task rules to automatically generate demonstrations without human teleoperation across diverse objects, environments, and randomized initial conditions. The generated data include basic skills such as grasping and placing, as well as long-horizon tasks that require both navigation and manipulation. We further introduce automated event labeling and trajectory categorization to filter demonstrations. These mechanisms enable fine-grained analysis of task progress, behavioral outcomes, and safety-related failures. Finally, we evaluate a range of imitation learning and VLA baselines across different task settings, revealing their performance characteristics and failure modes. Together, AeroManip-VLA enables scalable aerial manipulation data generation, structured trajectory analysis, and systematic VLA evaluation in simulation prior to real-world deployment.
发表机构
- National University of Singapore(新加坡国立大学)
- Beijing Institute of Technology(北京理工大学)
机构由 AI 辅助整理,请以论文原文为准。