发表机构
School of Automation, Beijing Institute of Technology; School of Mechanical Engineering, Beijing Institute of Technology(北京理工大学自动化学院; 北京理工大学机械工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究语言条件下四旋翼飞行问题,提出AeroAct模型,利用预训练视频扩散Transformer,结合相关采集设备和程序获取数据,经实验验证该模型能提升四旋翼飞行的目标跟踪和物体搜索性能,且对应策略可在实体四旋翼上执行。
AI 中文摘要
语言条件四旋翼飞行要求策略将语义目标落地,预测自我运动的视觉后果,并输出在快速变化的第一人称视角下保持平滑且可动态执行的控制参考。现有方法存在局限性。本文提出AeroAct,这是首个用于四旋翼导航的以动作中心的世界-动作模型。它采用预训练视频扩散Transformer预测轨迹-动作块,训练用未来帧作监督,部署时直接解码动作。还构建了相关管道,引入低成本采集设备和自引导程序。闭环模拟和实际实验表明其能提升性能且基于该模型的策略可在物理四旋翼上执行。
英文摘要
Language-conditioned quadrotor flight requires a policy to ground semantic goals, anticipate the visual consequences of ego-motion, and output control references that remain smooth and dynamically executable under rapidly changing first-person views. Existing aerial vision-language navigation and vision-language-action methods commonly use discrete actions, high-level waypoints, or instantaneous velocity commands, which provide limited supervision about how flight actions change future observations. We present AeroAct, an action-centered world-action model (WAM) for quadrotor navigation. To the best of our knowledge, AeroAct is the first WAM instantiated and demonstrated for real-world aerial flight. The model adapts a pretrained video diffusion Transformer to predict local trajectory-action chunks from egocentric visual history, proprioception, and language. Future first-person frames are used during training as dense consequence supervision, while deployment directly decodes actions without generating future video. To obtain aligned visual, state, language, and dynamically feasible action data, we build a DiffAero-based pipeline with complementary Isaac Lab and 3D Gaussian splatting renderers. We further introduce a low-cost handheld collection device that couples camera observations with motion estimates to recreate flight-like egocentric trajectories, and a self-guidance procedure that improves temporal consistency across overlapping trajectory chunks. Closed-loop simulation and real-world experiments show that temporal visual context improves target tracking and object-search performance, and that WAM-based policies can be executed on a physical quadrotor.