排名置信序列:随时有效的排行榜
Rank Confidence Sequences:Anytime-valid Leaderboards
浏览论文内容
中文总结 AI 辅助
本文提出排名置信序列,一种随时有效的排行榜方法,通过结合赌e过程和闭合检验,在任意停止规则下同时为所有模型提供有效排名集,支持随时检查且节省计算。
中文摘要 AI 辅助
排行榜根据模型在基准测试项上的平均得分对模型进行排名,并且在评估仍在进行时会被反复查阅。现有的模型排名置信区间只有在预先选定项目数量后计算一次时,才能控制其错误率。如果随着结果到达而重新计算,并且一旦结果看起来具有决定性就停止评估,则其错误率会超过名义水平。随时有效的方法在所有样本量下同时保持其保证,因此在任何停止规则下都有效。它们适用于单个模型的准确性、一对模型以及可能最佳的模型集合。对于成对比较,它们也给出排名。但当所有模型在同一项目上得分时,没有方法给出排名,这使得它们的得分具有依赖性。我们构建了排名置信序列:对于每个模型,一组包含其真实排名的排名集合,同时适用于所有模型和所有时间,在选定的错误水平$\alpha$下,在有限样本中成立。该构建将赌e过程(每对有序模型一个)与对模型可能排序的闭合检验相结合。它允许模型在某个项目上的得分之间存在任意依赖性。该方法有两个优点。排行榜可以在每个项目之后进行检查,而不会增加其错误率。每个模型的评估可以在关于其问题得到回答后立即停止,从而节省计算量。当结果仅检查一次(在中期或之后)时,与固定样本方法相比,损失的功效很小。本文通过模拟和公共排行榜数据量化了这些优点。
英文摘要
Leaderboards rank models by their average scores on benchmark items, and they are consulted repeatedly while the evaluation is still running. Existing confidence intervals for a model's rank control their error rate only if they are computed once, after a number of items chosen in advance. If they are recomputed as results arrive, and the evaluation stops once they look decisive, their error rate exceeds its nominal level. Anytime-valid methods keep their guarantees at all sample sizes simultaneously and hence under any stopping rule. They exist for the accuracy of one model, for one pair of models and for the set of models that may be best. For pairwise battles they also give ranks. None gives ranks when all models are scored on the same items, which makes their scores dependent. We construct rank confidence sequences: for every model, a set of ranks that contains its true rank, simultaneously for all models and at all times, at a chosen error level $α$, in finite samples. The construction combines betting e-processes, one for each ordered pair of models, with closed testing over the possible orderings of the models. It allows any dependence between the models' scores on an item. The method has two advantages. A leaderboard can be inspected after every item without inflating its error rate. The evaluation of each model can stop as soon as the question asked about it is answered, which saves compute. When results are examined only once, halfway through or later, little power is lost relative to fixed-sample methods. The paper quantifies these advantages in simulations and on public leaderboard data.
发表机构
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。