战略多智能体学习用于足球中所有球员的可解释行动估值
Strategic Multi-Agent Learning for Interpretable Action Valuation of All Players in Football
浏览论文内容
中文总结 AI 辅助
本研究提出一种受马尔可夫完美均衡启发的多智能体行动估值框架,通过分解Q值并利用可扩展决策状态,在95场J1联赛数据上实现比独立基线更上下文敏感的球员行动排名。
中文摘要 AI 辅助
评估足球中球员行动的价值需要考虑22名球员之间的战略互动,包括无球跑动和防守站位。现有的基于强化学习的方法通常将决策聚合在团队层面或独立估计球员价值,导致球员之间的战略相互依赖性未能得到充分体现。本研究提出了一种受马尔可夫完美均衡(MPE)启发的行动估值框架,适用于所有球员。每次控球被建模为有限时域动态博弈,每名球员被表示为自主智能体,其策略取决于当前比赛状态。MPE被用作激励性解概念而非精确均衡。为了提高可解释性,我们使用可扩展决策状态(EDMS)并将Q值分解为后继特征基和线性奖励权重向量。价值基通过线性时序差分(TD)初始化后进行非线性细化来估计。利用来自95场J1联赛比赛的追踪和事件数据,我们将所提出的公式与独立强化学习基线进行比较。由于两种公式在不同目标空间中定义TD误差,TD均方误差(MSE)仅用于公式内部的一致性检验。在EDMS固定的情况下,独立基线在99.21%的评估无球状态中将最高价值赋予向前移动,而所提出公式下最频繁的方向占17.63%。团队层面的平均Q值与基线赛季水平预期进球呈负相关,而与所提出公式呈弱正相关。定性分析说明了无球跑动和防守站位的上下文相关估值。总体而言,所提出的公式产生了更具上下文敏感性的行动排名,尽管比较并未隔离MPE启发的组成部分。
英文摘要
Valuing player actions in football requires accounting for strategic interactions among 22 players, including off-ball movements and defensive positioning. Existing reinforcement-learning-based methods commonly aggregate decisions at the team level or estimate player values independently, leaving strategic interdependence among players insufficiently represented. This study proposes an action valuation framework inspired by Markov perfect equilibrium (MPE) for all players. Each possession is modeled as a finite-horizon dynamic game, with each player represented as an autonomous agent whose policy depends on the current game state. MPE is used as a motivating solution concept rather than an exact equilibrium. To improve interpretability, we use Expandable Decision-Making States (EDMS) and decompose the Q-value into a successor-feature basis and a linear reward-weight vector. The value basis is estimated by linear TD initialization followed by nonlinear refinement. Using tracking and event data from 95 J1 League matches, we compare the proposed formulation with an independent reinforcement learning baseline. Because the two formulations define TD errors in different target spaces, TD MSE is used only for within-formulation consistency. With EDMS fixed, the independent baseline assigns the highest value to forward movement in 99.21% of evaluated off-ball states, whereas the most frequent direction under the proposed formulation accounts for 17.63%. Team-level average Q-values show a negative association with season-level expected goals for the baseline and a weakly positive association for the proposed formulation. Qualitative analyses illustrate context-dependent valuations of off-ball movements and defensive positioning. Overall, the proposed formulation produces more context-sensitive action rankings, although the comparison does not isolate the MPE-inspired component.
发表机构
- Nagoya University(名古屋大学)
- The University of Tokyo(东京大学)
- The Hong Kong University of Science and Technology(香港科技大学)
- RIKEN(理化学研究所)
机构由 AI 辅助整理,请以论文原文为准。