发表机构
Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出ShuttleArena羽毛球自博弈环境,采用PPO训练策略,实现了可解释的战术探测与竞争力提升,为交互式数字娱乐AI提供了有用测试平台。
AI 中文摘要
羽毛球是游戏AI领域中紧凑但极具挑战性的场景:玩家必须选择符合物理规律的羽毛球轨迹、预判对手的拦截动作,并恢复到某个场地位置,该位置的价值取决于对手的下一次回应。核心挑战在于击球选择与恢复动作并非可分离:最佳的恢复动作取决于击球引发的对手回应,而击球的价值则取决于击球者能否覆盖对手的回球。本文提出ShuttleArena,这是一个基于物理的羽毛球单打自博弈环境,融合了连续的羽毛球飞行、玩家拦截、结构化击球生成以及击球后恢复动作。该策略采用角色条件输出:在接发方回合采用掩码式拦截选择,在击球方回合采用因子化的击球动作,涵盖击球方位角、击球仰角、击球速度以及恢复目标,支持可解释的战术探测。训练采用近端策略优化(PPO),针对分阶段的检查点对手池进行自博弈,使用稀疏的回合终止结果奖励以及因子特定的恢复更新。通过冻结检查点博弈、受控战术探测、恢复动作消融实验、定性rollout以及人类数据合理性检验的评估显示,该方法在实现竞争力提升的同时,还能呈现出与对手相关的可解释击球几何特征与恢复行为变化。学习到的策略既产生了可识别的类羽毛球结构,也反映了模拟器的抽象特性,恢复干预实验表明,学习到的恢复行为对竞争力至关重要。这些结果表明,基于物理的隔网对抗运动是交互式数字娱乐AI的有用测试平台,因为它们要求智能体协调执行、站位以及与对手相对的战术价值。
英文摘要
Badminton is a compact but challenging domain for game AI: a player must choose a physically feasible shuttle trajectory, anticipate the opponent's interception, and recover to a court position whose value depends on the opponent's next response. The central challenge is that shot selection and recovery are not separable: the best recovery depends on the shot-induced opponent response, while the value of the shot depends on whether the hitter can cover the reply. This paper presents ShuttleArena, a physics-based singles badminton self-play environment that couples continuous shuttle flight, player interception, structured shot generation, and post-shot recovery. The policy uses role-conditioned outputs: a masked interception choice on receiver turns and a factorized hitter action over shot azimuth, shot elevation, shot speed, and recovery target, enabling interpretable tactical probes. Episodes are single rallies rather than full scored games, and training uses Proximal Policy Optimization (PPO) self-play against a staged checkpoint opponent pool with sparse terminal rally-outcome rewards and a factor-specific recovery update. Evaluation with frozen checkpoint play, controlled tactical probes, recovery ablations, qualitative rollouts, and a human-data sanity check shows competitive improvement together with interpretable opponent-conditioned changes in shot geometry and recovery behavior. The learned policies produce recognizable badminton-like structure while also reflecting the abstractions of the simulator, and the recovery intervention shows that learned recovery behavior is competitively important. These results suggest that physics-based racket sports are a useful testbed for interactive digital entertainment AI because they require agents to coordinate execution, positioning, and opponent-relative tactical value.