发表机构
Universitat de Barcelona; Universitat de Barcelona Institute of Complex Systems (UBICS); King Shared Services S.L; Dribia Data Research, S. L.(巴塞罗那大学; 巴塞罗那大学复杂系统研究所; 金共享服务有限公司; 德里比亚数据研究有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
我们提出一种无分布假设的 Solo 评分,将任务结果映射为有界得分并聚合以评定个体技能,在《糖果传奇》等游戏中以高置信度识别高、低技能玩家,且分类长期稳定。
AI 中文摘要
我们提出一种无分布假设的度量方法,用于在“玩家对环境”场景中评定个体技能水平,此类场景中参与者面对异质任务且没有直接对手。这种场景常见于数字平台、游戏、教育、金融以及人工智能智能体的基准评测中。它们兼具高随机性、难度差异极大的任务以及个体间未知的异质性。我们的度量将每个任务结果映射为一个有界性能得分,其总体均值为零,方差以 $1/3$ 为界,且不依赖于结果分布,从而使不同任务的得分可直接比较。将这些得分跨任务聚合,即为每个个体产生一个可解释的技能得分(即“Solo”评分),同时一个重排零模型可检验该得分是否超出纯偶然所能产生的水平。我们在手机游戏《糖果传奇》和《泡泡女巫3》中 $2\ imes10^5$ 名玩家的进度数据上验证了该方法。该度量能以高统计置信度识别高技能与低技能玩家,其分类在后续数百个关卡中保持稳定,并且得分的滑动窗口版本可追踪玩家在进度中的表现变化。
英文摘要
We propose a distribution-free metric to rate individual skill in ``player-versus-environment'' settings, where participants face heterogeneous tasks without direct opponents. Such settings are common in digital platforms, games, education, finance, and the benchmark evaluation of AI agents. They combine high randomness, tasks of widely varying difficulty, and unknown heterogeneity across individuals. Our metric maps each task outcome to a bounded performance score with zero population mean and variance bounded by $1/3$, regardless of the outcome distribution, so that scores are directly comparable across tasks. Aggregating these scores over tasks yields an interpretable skill score for each individual (the ``Solo'' rating), and a reshuffling null model tests whether that score exceeds what chance alone would produce. We validate the approach on the progression of $2\times10^5$ players in the mobile games Candy Crush Saga and Bubble Witch 3 Saga. The metric identifies high- and low-skill players with high statistical confidence, their classification persists over hundreds of subsequent levels, and a windowed version of the score tracks changes in performance along progression.