arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种用于钝体尾流主动流动控制的两阶段基于模型的强化学习方法

A Two-Stage, Model-Based Reinforcement Learning Approach for Active Flow Control of Bluff Body Wakes

Aayushman Sharma, Suman Chakravorty

arXiv 2609.08436首次发表:更新:

AI 中文总结

本文提出一种两阶段基于模型的强化学习方法,利用稀疏表面压力观测实现钝体尾流主动控制,在Re=100圆柱尾流模拟中完全抑制涡脱落并降低44%阻力。

AI 中文摘要

本文针对具有未知且不稳定平衡点的高维非线性系统,利用稀疏部分观测,提出了一种数据驱动的输出反馈方法,用于解决无限时域最优控制问题。该方法基于无限时域问题的转移加调节分解:一个有限时域非线性转移将系统驱动到动力学可由关于未知工作点的线性模型良好近似的区域,然后在该区域内辨识的无限时域线性调节器完成稳定化。我们将该框架扩展到部分观测设置,通过结合基于ARMA的信息状态构造与两阶段控制架构:在信息状态上采用迭代线性二次调节器(iLQR)方法将系统驱动到平衡邻域,该平衡邻域在无目标先验知识的情况下隐式发现,而局部辨识的时不变ARMA模型提供无限时域调节器以实现渐近稳定。该方法无需伴随求解器、降阶模型或全状态访问。我们在高保真纳维-斯托克斯模拟中,于雷诺数Re=100的圆柱尾流上验证了该方法,仅使用八个表面压力传感器,比近期基于模型的强化学习方法少一个数量级。该控制器实现了对涡脱落引起的升力振荡的完全抑制,并且相对于未受控基线,总阻力降低了44%。

英文摘要

This paper develops a data-driven, output-feedback approach to the infinite-horizon optimal control of high-dimensional nonlinear systems with unknown and unstable equilibria, using sparse partial observations. The approach builds on the transfer-plus-regulation decomposition of the infinite-horizon problem: a finite-horizon nonlinear transfer drives the system into a region where the dynamics are well-approximated by a linear model about the unknown operating point, and an infinite-horizon linear regulator identified within that region completes stabilization. We extend this framework to the partially observed setting by combining an ARMA-based information-state construction with a two-stage control architecture: an iterative linear quadratic regulator (iLQR) approach on the information state drives the system to the equilibrium neighborhood, discovered implicitly without prior knowledge of the target, and a locally identified time-invariant ARMA model provides the infinite-horizon regulator for asymptotic stabilization. The method requires no adjoint solver, reduced-order model, or full-state access. We validate the approach on high-fidelity Navier-Stokes simulations of the cylinder wake at $\mathrm{Re}=100$ using only eight surface pressure sensors, an order of magnitude fewer than recent model-based RL methods. The controller achieves complete suppression of vortex-shedding-induced lift oscillations and a $44\%$ reduction in total drag relative to the uncontrolled baseline.

Comments9 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑