arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FlowPilot:面向敏捷无人机导航的实时世界-动作建模

FlowPilot: Real-Time World-Action Modeling for Agile UAV Navigation

Runqing Wang, Ding Yu, Pengyuan Min, Xinhong Zhang, Wei Xiao, Yu Hu, Jie Chen, Fu Zhang, Gang Wang

arXiv 2608.00635首次发表:更新:

发表机构

Beijing Institute of Technology; Zhongguancun Academy; Shanghai Jiao Tong University; The University of Hong Kong(北京理工大学; 中关村学院; 上海交通大学; 香港大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FlowPilot是一种基于流匹配的紧凑世界-动作模型,采用双流混合Transformer,在三级深度金字塔数据上训练,可实现敏捷无人机实时导航,性能优于基线,能在杂乱环境中以5.5m/s速度运行。

AI 中文摘要

我们提出FlowPilot,一种用于基于深度的实时机载无人机导航的紧凑世界-动作模型。与需要局部重建的“先建图再优化”流水线或缺乏显式场景预测的端到端策略不同,FlowPilot通过流匹配联合对未来深度观测和可执行轨迹进行去噪。双流混合Transformer通过共享注意力耦合视频与动作专家,使未来场景预测和轨迹生成能相互赋能。部署时,模型以动作为中心运行,仅输出轨迹。为确保可跟踪性,动作被参数化为7阶伯恩斯坦多项式:当前状态约束初始控制点,网络预测5个自由控制点,生成具有闭式速度、加速度和加加速度的C²连续参考。FlowPilot在覆盖高通量仿真、光致真仿真和真实机载数据的三级深度金字塔上训练。在闭环仿真中,它在障碍物增多和指令速度达8m/s的情况下,优于基于学习和优化的基线。在物理四旋翼上,完整感知-动作流水线在Jetson Orin NX上运行耗时不足18ms,仅使用机载感知与计算,在杂乱室内和森林环境中速度达5.5m/s。

英文摘要

We present FlowPilot, a compact world-action model for real-time onboard UAV navigation from depth. Unlike map-then-optimize pipelines that require local reconstruction or end-to-end policies that lack explicit scene prediction, FlowPilot jointly denoises future depth observations and executable trajectories with flow matching. A dual-stream mixture-of-transformers couples video and action experts through shared attention, allowing future-scene prediction and trajectory generation to inform each other. At deployment, the model runs action-centrically and outputs only a trajectory. To ensure trackability, actions are parameterized as degree-7 Bernstein polynomials: the current state constrains the initial control points, and the network predicts five free control points, yielding C^2-continuous references with closed-form velocity, acceleration and jerk. FlowPilot is trained on a three-level depth pyramid spanning high-throughput simulation, photorealistic simulation, and real onboard data. In closed-loop simulation, it outperforms learning- and optimization-based baselines under increasing clutter and commanded speeds up to 8m/s. On a physical quadrotor, the full perception-to-action pipeline runs in under 18ms on a Jetson Orin NX and reaches 5.5m/s in cluttered indoor and forest environments using only onboard sensing and computation.

Comments8 pages, 9 figures, 2 tables, submitted to IEEE Robotics and Automation Letters (RA-L)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑