arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

引擎平等,人类不平等:引擎评估的平等国际象棋局面中可重复的结果偏差

Engine-Equal, Human-Unequal: A Reproducible Outcome Skew in Engine-Assessed Equal Chess Positions

Jesung Park

arXiv 2607.25655首次发表:更新:

AI 中文总结

研究国际象棋中引擎判定平等的局面,人类结果却不平衡,存在结果偏差。通过特定方法测量偏差,发现其可在多种划分中重现,典型偏差小但稳定,即便评估有信心也非人类结果充分统计量,因果问题待随机对照研究。

AI 中文摘要

在强大的引擎判定基本平等(Stockfish 18评估在零的10个兵分值内且深度稳定)且人类在Lichess上实际达成的国际象棋开局局面中(2025年10月;1661个局面,1610万次出现),人类结果并不平衡。局面存在结果偏差,即其实际结果与玩家等级分预测结果之间的差距,偏差方向是自然达成局面的稳定属性:一些局面利于白方,另一些利于黑方。这些偏差在三个重新划分中重现——不相交的玩家账户集(主要)、时间和不相交的等级分区间——以及在八个月后的样本外月份。在主要划分中,每个局面的偏差在每个账户组中测量一次,复制斜率询问在去除等级分和开局家族效应后一个测量值对另一个的预测程度:1表示无衰减的延续;0表示无线性关系。我们发现斜率为0.69(家族聚类95%置信区间[0.65, 0.74]),在最流行、测量最佳的局面上升至0.94。斜率值取决于局面组合。偏差的存在是不变的论断:它在我们测试的每个更严格的评估区间、搜索深度、校准和流行度截止条件下都存在,并分别在快棋和超快棋中重现。典型偏差较小(中位数$|\delta| \approx 0.018$,约为白方得分的两个百分点),但它逐局面地在不相交账户中重现。在这些局面上,处于劣势的一方思考时间也更长。即使评估最有信心,它也不是人类结果的充分统计量。结果是观察性的,因果问题留待预先注册的随机对照研究解决。

英文摘要

Among chess opening positions that a strong engine judges essentially equal (Stockfish 18 evaluation within 10 centipawns of zero, depth-stable) and that humans actually reach on Lichess (October 2025; 1,661 positions, 16.1M occurrences), human results are not balanced. Positions carry outcome skews, each the gap between its games' actual results and what the players' ratings predict, whose directions are stable properties of the naturally-reached position: some positions favour White, others Black. These skews reproduce across three re-partitions -- disjoint player-account sets (primary), time, and disjoint rating bands -- and on an out-of-sample month eight months later. On the primary split, each position's skew is measured once in each account group, and the replication slope asks how well one measurement predicts the other after removing rating and opening-family effects: one means undiminished carry-over; zero, no linear relation. We find 0.69 (family-clustered 95% CI [0.65, 0.74]), rising to 0.94 on the most-popular, best-measured positions. The slope's value depends on the position mix. Existence is the invariant claim: it survives every tighter evaluation band, search depth, calibration, and popularity cutoff we test, and replicates within blitz and rapid separately. The typical skew is small (median $|δ| \approx 0.018$, about two percentage points of White score), yet it reproduces, position by position, across disjoint accounts. At these positions the disfavoured side also thinks longer. Even where the evaluation is most confident, it is not a sufficient statistic for human outcomes. The result is observational, and the causal question is left to a pre-registered randomised companion study.

Comments27 pages, 3 figures, 3 tables. Code and data: https://doi.org/10.5281/zenodo.21629354

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑