达到标准杆:面向可证明最优四边形块分解的强化学习
Playing to Par: Reinforcement Learning for Provably Optimal Quadrilateral Block Decompositions
浏览论文内容
中文总结 AI 辅助
本文提出用强化学习智能体直接操作网格半边结构,通过局部编辑和基于网格连通性的卷积策略,生成达到离散高斯-博内下界(标准杆)的四边形分解,在保留域上显著优于Gmsh,实现可证明最优的网格生成。
中文摘要 AI 辅助
平面域的四边形块分解通过其是否完整、元素形状是否良好以及顶点不规则数量来评判。最后一个性质存在可证明的下界:离散高斯-博内恒等式仅基于给定域的全部四边形网格的拓扑和角点角度,强制规定了总顶点不规则度的下限。我们训练一个强化学习智能体来构建达到此下界(称为“标准杆”)的分解。该智能体通过局部编辑直接作用于网格的半边数据结构,其策略网络的卷积遵循网格自身的连通性,因此可以不加修改地应用于比训练时所见更大的域。奖励直接针对该下界,且是稀疏的:在边数超过八的域上,随机策略无法达到该下界。我们通过在易于构造的最优网格上进行行为克隆,将其反向生成为演示,再使用PPO进行训练,从而克服了这一探索障碍。在96个保留域上,该智能体在每个域上都生成了全四边形网格,平均95.7个可用,90个可证明最优;而Gmsh在相同元素数量下最强的配置完成了51个,可用38个,最优0个,即使元素数量增至三到十四倍,也从未产生更规则的网格。在64个大小为训练规模两倍的域上,该智能体全部完成,62个可用,且与相同元素数量下Gmsh的39相比,其超出标准杆的中位数余量低于1。
英文摘要
A quadrilateral block decomposition of a planar domain is judged by whether it is complete, whether its elements are well shaped, and how many of its vertices are irregular. The last has a provable floor: the discrete Gauss-Bonnet identity enforces a lower bound on the total vertex irregularity of any all-quadrilateral mesh of a given domain purely based on its topology and corner angles. We train a reinforcement learning agent to build decompositions that reach this bound, which we call par. It acts directly on the mesh's half-edge data structure through local edits, with a policy network whose convolutions follow the mesh's own connectivity, so it applies unchanged to domains larger than any seen in training. The reward targets the floor directly, and it is sparse: random play reaches it on no domain with more than eight sides. We overcome this exploration barrier via behaviour cloning on optimal meshes that are trivial to construct, walked backward into demonstrations, before training it with PPO. On 96 held-out domains the agent produces an all-quadrilateral mesh on every one, a usable one on 95.7 on average, and a provably optimal one on 90; Gmsh's strongest configuration at the same element count completes 51, is usable on 38 and optimal on none, and even at three to fourteen times the elements never produces a more regular mesh. On 64 domains twice the training size the agent completes all, is usable on 62, and keeps a median excess over par below one against Gmsh's 39 at the same element count.