arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AeroDPO:释放具备高保真感知与自动偏好优化的轻量级无人机导航能力

AeroDPO: Unleashing Lightweight UAV Navigation with High-Fidelity Perception and Automated Preference Optimization

Peng Xu, Chengcheng Wang, Shaohua Wan

arXiv 2608.07557首次发表:更新:

AI 中文总结

本文提出AeroDPO,通过高保真感知与自动偏好优化,使20亿参数轻量级无人机导航模型达到70亿参数模型的成功率,同时降低分布外场景碰撞率,成为自主空中智能体新SOTA。

AI 中文摘要

无人机视觉语言导航(UAV-VLN)需在复杂三维环境中实现快速反应式控制。近期极简端到端范式展现出巨大潜力,但通常依赖包含数十亿参数的大规模语言模型,在实际边缘部署中会产生难以接受的延迟。本文挑战这种对高参数模型的依赖,全面跨尺度评估揭示出关键洞见:感知质量根本上优于语言推理能力。我们证明,配备高保真视觉输入的20亿参数轻量级模型,其总体成功率完全匹配70亿参数大规模基线模型。然而,这种极简策略暴露出纯行为克隆(BC)固有的基本鲁棒性缺陷:缺乏显式负反馈,智能体无法内化鲁棒空间约束,在分布外(OOD)场景中表现出惊人的碰撞率。为在不依赖不可扩展人工标注的情况下克服这一漏洞,我们提出AeroDPO——一种由确定性物理模拟状态回退驱动的零成本自动直接偏好优化(DPO)流程。检测到碰撞时,系统自动回退环境以提取因果推理错误作为被拒绝动作,应用解耦特权干预合成避障偏好动作,并利用离线视觉语言检查器过滤视觉歧义。通过为20亿参数模型配备此自动数据飞轮,AeroDPO在未映射场景上将成功率提升至49.16%,同时大幅降低碰撞率,为自主智能体建立新的SOTA。

英文摘要

Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid and reactive control in complex 3D environments. Recent minimalist end-to-end paradigms show great promise but typically rely on massive language models containing billions of parameters, incurring prohibitive latency for real-world edge deployment. In this paper, we challenge this parameter-heavy reliance. Comprehensive cross-scale evaluations reveal the critical insight that perception quality fundamentally outweighs language reasoning capacity. We demonstrate that a lightweight 2B model equipped with high-fidelity visual inputs completely matches the overall success rates of massive 7B baselines. However, this minimalist policy exposes a fundamental robustness flaw inherent to pure Behavior Cloning (BC). Lacking explicit negative feedback, the agent fails to internalize robust spatial constraints and exhibits alarming collision rates in out-of-distribution (OOD) scenarios. To overcome this vulnerability without relying on unscalable human annotations, we propose AeroDPO, a zero-cost automated Direct Preference Optimization pipeline driven by deterministic physical simulation state rollback. Upon detecting collisions, the system autonomously rewinds the environment to extract causal reasoning errors as rejected actions, applies decoupled privileged interventions to synthesize collision-avoidance preferred maneuvers, and leverages an offline vision language inspector to filter visual ambiguities. By equipping our 2B model with this automated data flywheel, AeroDPO boosts success rates to 49.16% on unmapped scenarios while drastically suppressing collision rates, establishing a new SOTA for autonomous aerial agents.

Comments7 pages, 3 figures, 4 tables, Code is available at: [https://github.com/XuPeng23/AeroDPO]

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑