物理约束软Actor-评论家算法用于模拟器在环的油藏历史拟合
Physics-Constrained Soft Actor-Critic for Simulator-in-the-Loop Petroleum Reservoir History Matching
浏览论文内容
中文总结 AI 辅助
该研究针对仅提供昂贵模拟器的科学校准问题,提出物理约束SAC算法,在PUNQ-S3数据集上实现油藏历史拟合的高精度匹配,验证了基于强化学习的昂贵科学模拟器校准的可行性。
中文摘要 AI 辅助
许多科学校准问题仅提供昂贵的可执行模拟器,无法获取梯度且大规模训练数据生成不切实际。本文将油藏历史拟合作为该类AI问题的实例,将其建模为物理约束、模拟器在环的策略搜索问题。我们将CMG IMEX全物理模拟器封装为Gymnasium环境,采用软Actor-评论家(SAC)算法学习连续孔隙度、定向渗透率及井表皮参数的随机提议分布。每次交互生成并执行油藏案例,使模拟与观测生产响应对齐,返回的奖励结合多响应失配与物理无效属性的惩罚。离策略回放复用昂贵的模拟器反馈,最大熵学习保留探索性。与正向代理方法不同,该策略学习评估位置而非替代模拟器,所有保留的候选均经IMEX验证。在PUNQ-S3上200次调用预算下,最佳有效候选达到类别宏平均归一化均方误差(NMSE)0.0285、决定系数R²=0.9324,有界匹配得分为97.23%。井级分析进一步揭示了聚合指标掩盖的局部水率失效。这些结果为基于强化学习的昂贵科学模拟器校准建立了全物理概念验证。
英文摘要
Many scientific calibration problems expose only an expensive executable simulator, making gradients unavailable and large-scale training-data generation impractical. We study petroleum reservoir history matching as an instance of this broader AI problem and formulate it as physics-constrained, simulator-in-the-loop policy search. Our method wraps the CMG IMEX full-physics simulator as a Gymnasium environment and uses Soft Actor-Critic (SAC) to learn a stochastic proposal distribution over continuous porosity, directional-permeability, and well-skin parameters. Each interaction generates and executes a reservoir case, aligns simulated and observed production responses, and returns a reward that combines multi-response mismatch with penalties for physically invalid properties. Off-policy replay reuses costly simulator feedback, while maximum-entropy learning preserves exploration. Unlike forward-surrogate approaches, the policy learns where to evaluate rather than learning to replace the simulator; every retained candidate is validated by IMEX. Under a 200-call budget on PUNQ-S3, the best valid candidate achieves category-macro NMSE $0.0285$, $R^2=0.9324$, and a bounded match score of $97.23\%$. Well-level analysis further exposes localized water-rate failures hidden by pooled metrics. These results establish a full-physics proof of concept for reinforcement-learning-based calibration of expensive scientific simulators.