发表机构
ETH Zurich; Belimo Automation AG; University of Padova(苏黎世联邦理工学院; 贝利莫自动化股份公司; 帕多瓦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对冰壶战术决策的建模挑战,采用适配有限时间范围结构的DDPG算法,通过自监督学习获得的智能体在简化四石冰壶中可匹敌专家启发式策略,其评论家还能提供连续动作空间的价值估计以支持战术分析。
AI 中文摘要
冰壶因决策过程的战术复杂性常被称为“冰上国际象棋”。然而与国际象棋不同,冰壶在机器学习视角下仍未得到充分探索,现有研究主要局限于统计方法。我们提出一种强化学习框架,能够定量评估和比较冰壶中的战术选项。该游戏存在若干建模挑战:连续的状态与动作空间、反映选手技能差异的随机动作结果,以及执行动作的微小扰动对状态转换高度敏感。为解决这些问题,我们采用深度确定性策略梯度(Deep Deterministic Policy Gradient,DDPG)演员-评论家算法,该算法适配并利用了游戏的有限时间范围结构。我们的实验表明,无需任何人工标注数据,即可通过完全自监督的方式习得有效的冰壶策略:在简化的四石变体中,当手工设计的专家启发式策略接近最优时,所学习的智能体与该策略表现相当,我们针对该变体固有的先手优势(hammer advantage)量化了这种对等性。除了习得的策略外,所学习的评论家还能在整个连续动作空间上提供密集的价值估计,支持战术选项的定量比较,可应用于赛后表现分析、运动员备战期间的决策支持等场景。
英文摘要
Curling is often referred to as "Chess on Ice", owing to the tactical complexity of its decision-making process. Yet unlike chess, curling remains largely underexplored from a machine learning perspective, with prior work confined mainly to statistical approaches. We propose a reinforcement learning framework capable of quantitatively evaluating and comparing tactical options in curling. The game poses several modeling challenges: continuous state and action spaces, stochastic action outcomes reflecting player skill variability, and state transitions that are highly sensitive to small perturbations in the executed action. To address them, we employ the Deep Deterministic Policy Gradient actor-critic algorithm, adapted to exploit the finite-horizon structure of the game. Our experiments show that effective curling strategies can be acquired in a fully self-supervised manner, without any human-annotated data: on a reduced four-rock variant, the learned agent matches a hand-crafted expert heuristic in a regime where that heuristic is close to optimal, a parity we quantify against the intrinsic hammer advantage of the variant. Beyond the resulting policy, the learned critic provides a dense value estimate over the entire continuous action space, enabling the quantitative comparison of tactical alternatives for applications such as post-game performance analysis and decision support during athlete preparation.
Comments10 pages, 8 figures