arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于视觉的敏捷间隙穿越学习:带暖启动评论家的可微分仿真

Learning Vision-Based Agile Gap Traversal: Differentiable Simulation with a Warm-Started Critic

Nuthasith Gerdpratoom, Tianchen Sun, Yichao Gao, Lin Zhao

arXiv 2609.30696首次发表:更新:

发表机构

National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出两阶段强化学习框架,结合可微分仿真QPG与评论家暖启动,高效训练视觉间隙穿越策略,提升训练效率与成功率,并泛化至新形状和平台。

AI 中文摘要

穿越狭窄间隙对自主四旋翼飞行器而言极具挑战性,尤其是当控制指令直接来自高维视觉观测时。现有的端到端方法通常依赖行为克隆或通过可微分仿真进行全轨迹时间反向传播(BPTT),这可能会限制策略性能或导致高昂的训练成本。我们提出了一种两阶段强化学习框架,用于更高效地训练以自我为中心的视觉运动间隙穿越策略,该框架利用基于可微分仿真的准解析策略梯度(QPG)和评论家暖启动。该框架采用QPG以避免通过视觉渲染进行反向传播,从而降低计算和内存成本,同时提高样本效率。在第一阶段,使用包括间隙几何在内的特权观测训练专家演员和评论家。与先前的间隙穿越方法不同,我们利用QPG的训练不需要将智能体重置到优化的参考轨迹上。在第二阶段,使用来自两个自我中心相机的二值间隙掩码和低维观测训练视觉策略,而其特权评论家则从第一阶段暖启动。与冷启动评论家或使用全轨迹BPTT相比,这显著提高了训练效率和穿越成功率。我们的框架在系统参数变化时无需重新训练专家演员,从而比基于动作监督的最先进视觉间隙穿越方法在无人机平台间实现更高效的泛化。学习到的视觉策略还能泛化到形状未见过的间隙。大量的真实世界实验进一步证明了使用在线渲染的二值掩码进行稳健的间隙穿越。除了间隙穿越,所提出的框架具有通用性,可扩展到其他视觉运动机器人学习任务。

英文摘要

Traversing narrow gaps is challenging for autonomous quadrotors, especially when control commands come directly from high-dimensional visual observations. Existing end-to-end methods often rely on behavior cloning or full-rollout backpropagation through time (BPTT) via differentiable simulation, which can limit policy performance or incur high training costs. We propose a two-stage reinforcement learning framework for more efficient ego-centric visuomotor gap-traversal policy training, leveraging quasi-analytical policy gradients (QPG) via differentiable simulation and critic warm-starting. The framework utilizes QPG to avoid backpropagation through visual rendering, reducing computation and memory costs while improving sample efficiency. In the first stage, an expert actor and critic are trained using privileged observations, including gap geometry. Unlike prior gap-traversal approaches, our training utilizing QPG does not require resetting the agent along optimized reference trajectories. In the second stage, a visual policy is trained using binary gap masks from two ego-centric cameras and low-dimensional observations, while its privileged critic is warm-started from the first stage. This substantially improves training efficiency and traversal success compared with cold-starting the critic or using full-rollout BPTT. Our framework does not require retraining the expert actor when system parameters change, enabling more efficient generalization across drone platforms than state-of-the-art visual gap-traversal methods based on action supervision. The learned visual policy also generalizes to gaps with unseen shapes. Extensive real-world experiments further demonstrate robust gap traversal using binary masks rendered online. Beyond gap traversal, the proposed framework is generic and can be extended to other visuomotor robot learning tasks.

Comments8 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑