arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13553cs.ROcs.LGcs.SYeess.SY

通过强化学习在非定常流中进行流量感知最优导航

Flow-aware Optimal Navigation in Unsteady Flows through Reinforcement Learning

  • Politecnico di Milano(米兰理工大学)
  • MOX(数学优化与工程计算实验室)

机构由 AI 辅助整理,请以论文原文为准。

Andrea Maria Braghin, Nicolò Botteghi, Matteo Tomasetto, Andrea Manzoni, Gabriele Cazzulani

AI总结:

研究在非定常流中自主机器人导航问题,用TD3算法训练智能体在双涡旋流中到达目标。评估五种生物启发观测策略及全局流参数影响,发现速度感知与涡度传感器效用有权衡,明确参数会降性能,为机器人导航从模拟到现实提供见解。

AI中文摘要:

由于部分可观测性和现实环境的不可预测性,在非定常时变流体流中进行自主机器人导航仍然是一个基本挑战。机器人中使用的经典最优控制框架需要不切实际的先验全局流知识,而生物系统能够通过利用局部感官线索成功导航。本文提出一种使用TD3算法的强化学习方法,训练自主智能体在参数化混沌双涡旋流中到达任意目标。为研究最优感官机制,评估了基于相对位置、局部速度或局部涡度测量以及短期记忆变体的五种受生物启发的观测策略。此外分析了为智能体提供明确全局流参数的影响。数值结果表明,能够感知并记住一定数量流速测量的智能体性能最高。实验揭示了传感器效用的权衡:速度感知智能体优化能源效率,而涡度传感器提供更好的结构映射并实现更好的目标接近度。纳入明确全局流参数会降低导航性能。这表明基于强化学习的自主系统在限于隐式流表示时会制定更稳健和通用的策略。研究结果为改善受生物启发的机器人导航从模拟到现实环境的转变提供了见解。

英文摘要:

Autonomous robotic navigation in nonstationary time-varying fluid flows remains a fundamental challenge due to partial observability and the unpredictability of realistic environments. While classical optimal control frameworks employed in robotics require unrealistic a-priori global flow knowledge, biological systems are able to navigate successfully by exploiting localized sensory cues. In this work we present a reinforcement learning approach using the TD3 algorithm to train autonomous agents to reach arbitrary targets within a parametric, chaotic double-gyre flow. To investigate optimal sensory mechanisms, we evaluate five bio-inspired observation strategies based on relative position, local velocity or local vorticity measures, and short-term memory variants. Additionally, we analyze the impact of providing agents with explicit global flow parameters. Numerical results demonstrate that an agent that is able to sense and remember a set number of flow velocity measures achieves the highest performance. The experiments reveal a trade-off in sensor utility: velocity-aware agents optimize energy efficiency, whereas vorticity sensors provide superior structural mapping and achieve better target proximity. Incorporating explicit global flow parameters is shown to decrease navigation performance. This behavior suggests that reinforcement learning-based autonomous systems develop more robust and general policies when restricted to implicit flow representations. The presented results offer insights for improving the transition of bio-inspired robotic navigation from simulation to real-world environments.

↑