arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2606.03963cs.ROcs.AI

面向视觉条件的无人机导航的自优化智能体强化学习

AgenticRL: Agentic Reinforcement Learning with Self-Refinement for Complex UAV Navigation

  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Roohan Ahmed Khan, Yasheerah Yaqoot, Amir Atef Habel, Muhammad Ahsan Mustafa, Dzmitry Tsetserukou

AI总结:

提出AgenticRL框架,利用多模态GPT智能体自动设计奖励函数、通过闭环自改进优化策略,在多种无人机导航任务中提升性能并实现高成功率。

AI中文摘要:

深度强化学习在使自主机器人学习复杂导航任务方面显示出巨大潜力。然而,其实际应用仍然严重依赖于人工设计的奖励函数和重复的手动微调,这既耗时又无法保证在目标任务中取得高成功率。本文提出了AgenticRL,一种智能体引导的强化学习框架,用于提高无人机导航任务中奖励设计、策略优化和实际部署的自主性。AgenticRL使用多模态生成预训练变换器(GPT)智能体来解释任务信息和视觉场景观察,生成特定于任务的奖励函数,使用近端策略优化(PPO)算法训练策略,然后通过诊断包评估训练后的策略作为批评者,生成反馈。基于该反馈,智能体识别失败模式并在闭环自改进过程中优化奖励函数。为了在推理期间进一步利用多模态GPT智能体,AgenticRL使用真实世界图像和自然语言任务信息自动识别活动场景并选择适当的训练策略执行。该框架在多种导航任务上进行了评估,包括穿越门、避障、穿越墙障并着陆、轨迹跟踪和运动行为学习。实验结果表明,与初始奖励相比,闭环优化过程将策略行为提升了71%。我们还展示了所提出框架的仿真到现实迁移,实现了91%的真实世界成功率和94%的仿真到现实准确率。

英文摘要:

Deep reinforcement learning enables autonomous robots to learn complex navigation tasks, but still relies heavily on time consuming manual reward design and fine tuning. Existing automated reward generation and refinement methods reduce this effort, yet often lack task-level behavioral diagnosis for directing subsequent reward revisions. We introduce AgenticRL, a multimodal closed loop framework in which role-specialized agents generate executable rewards, diagnose failures of the resulting policies, formulate targeted refinement instructions, and regenerate improved rewards. Before training, a task grounding stage automatically selects a compatible action profile, together with its observation and reward interfaces. Each generated reward is used to train a policy using Proximal Policy Optimization (PPO), which is subsequently evaluated under randomized conditions. Task-level behavioral, geometric, and safety measurements are organized into a structured diagnosis packet and jointly analyzed with the current reward code, task specification, behavioral summary, and visual scene context. Unlike one-shot reward generation, human-guided refinement, or broad candidate search, AgenticRL uses automated diagnosis of the behavior induced by a reward to direct its next revision. We evaluate the framework across eight UAV tasks covering navigation, obstacle interaction, trajectory tracking, agile manoeuvres, and cluttered flight. Under the reported comparative evaluation, AgenticRL achieves success rates of 100% in racing and 88% in cluttered navigation, exceeding the strongest Eureka and Text2Reward baselines, respectively. Reward refinement increases mean simulation success from 37.2% to 96.4%, while the resulting policies achieve a collective real-world success rate of 90.0% and a sim-to-real accuracy of 93.4%.

↑