发表机构
Mines Paris - PSL University Centre for Material Forming (CEMEF) CNRS(巴黎矿业 - 巴黎文理研究大学材料成型中心(CEMEF)法国国家科学研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究主动流动控制中控制器设计难题,提出用编码智能体直接搜索显式可执行反馈律的启发式学习范式及受限协议,经多基准评估,其启发式控制器性能优、可解释,是传统强化学习的可靠互补替代方案。
AI 中文摘要
主动流动控制涉及非线性动力学、部分观测和计算成本高昂的模拟,这使得控制器设计极具挑战性。深度强化学习(DRL)已成为解决此类问题的强大框架,但其成功通常依赖大量模拟器交互,且产生的神经网络策略决策过程往往难以解释。在这项工作中,我们研究了一种不同的范式:不是优化神经网络参数,而是使用现代编码智能体直接搜索显式可执行反馈律。我们引入了一种受限启发式学习协议,其中智能体通过公共基准接口迭代地提出、评估和修改控制器实现。在13个跨越一维、二维和三维问题的主动流动控制基准上对所提出的框架进行了评估,并在相同模拟预算下与最强的可用DRL基线进行了比较。发现的启发式控制器在13个环境中的10个中匹配或优于最佳DRL策略,同时保持紧凑、可解释且可直接检查。除了总体性能外,所得控制器揭示了具有物理意义的反馈机制,能成功转移到更具挑战性的配置中,并且在不同的雷诺数和瑞利数、致动器数量以及观测稀疏性下仍具有竞争力。这些结果表明,通过编码智能体进行启发式学习构成了传统强化学习的可靠且互补的替代方案,将具有竞争力的性能与可物理解释的控制器表示相结合。提示和源代码可在该https URL获取。
英文摘要
Active flow control involves nonlinear dynamics, partial observations, and computationally expensive simulations, making controller design particularly challenging. Deep reinforcement learning (DRL) has emerged as a powerful framework for such problems, but its success typically relies on large numbers of simulator interactions and produces neural-network policies whose decision process often remains difficult to interpret. In this work, we investigate a different paradigm: instead of optimizing neural-network parameters, we use modern coding agents to search directly for explicit executable feedback laws. We introduce a constrained heuristic-learning protocol in which an agent iteratively proposes, evaluates, and revises controller implementations while interacting exclusively through the public benchmark interface. The proposed framework is evaluated on 13 active flow-control benchmarks spanning one, two, and three-dimensional problems and compared against the strongest available DRL baselines under identical simulation budgets. The discovered heuristic controllers match or outperform the best DRL policy in 10 of the 13 environments while remaining compact, interpretable, and directly inspectable. Beyond aggregate performance, the resulting controllers reveal physically meaningful feedback mechanisms, transfer successfully across more challenging configurations, and remain competitive under varying Reynolds and Rayleigh numbers, actuator counts, and observation sparsity. These results suggest that heuristic learning through coding agents constitutes a credible and complementary alternative to conventional reinforcement learning, combining competitive performance with physically interpretable controller representations. Prompts and source code are available at https://github.com/DonsetPG/fluid-heuristic-learning.