AI 中文总结
本文提出 BB-EDGE 框架,将 LLM 排行榜建模为有向图,通过块因子化 e-过程与 e-Holm 方法实现任意时刻有效的成对优势认证和 FWER 控制,并在实验中验证其高效性。
AI 中文摘要
大语言模型(LLM)排行榜通过根据模型在固定基准上的平均表现对模型进行排名来比较模型能力。然而,不同运行之间评估结果的变异性可能产生关于模型在基准上优越性的无根据声明,而排行榜的更新会加剧这一风险。在本文中,我们提出 BB-EDGE(用于有向图评估的基准加权与块因子化 e-过程),这是一个原则性框架,将 LLM 排行榜表示为一个有向图,其边证明成对平均表现优势,并具有任意时刻有效的族系错误率(FWER)控制。具体而言,对于每个方向,BB-EDGE 通过将证据分解为协议定义的块,并根据相应块权重分配赌注,构建一个经验-伯恩斯坦 e-过程,然后在这些 e-过程上应用直接 e-Holm 方法,以将方向性优势认证为边。理论上,我们在异质基准平均零假设下刻画了权重比例线性赌注,并证明了在任意块内和跨对依赖下的任意时刻 FWER 控制。BB-EDGE 还支持任意时刻有效的 Top-k 认证和同时排名区间。在合成数据和四个真实世界基准上的大量实验表明,BB-EDGE 在保持任意时刻 FWER 控制的同时实现了高效率。
英文摘要
Large language model (LLM) leaderboards compare model capabilities by ranking models according to their mean performance on fixed benchmarks. However, variability in evaluation outcomes across runs may produce unsupported claims of model superiority on the benchmark, a risk compounded by leaderboard updates. In this paper, we propose BB-EDGE (Benchmark-Weighted and Block-Factorized e-processes for Directed Graph Evaluation), a principled framework that represents an LLM leaderboard as a directed graph whose edges certify pairwise mean-performance advantages, with anytime-valid family-wise error rate (FWER) control. Concretely, for each direction, BB-EDGE constructs an empirical-Bernstein e-process by factorizing evidence over protocol-defined blocks and assigning stakes proportional to the corresponding block weights, then applies direct e-Holm across these $e$-processes to certify directional advantages as edges. Theoretically, we characterize weight-proportional linear stakes under heterogeneous benchmark-average nulls and prove anytime FWER control under arbitrary within-block and cross-pair dependence. BB-EDGE further supports anytime-valid Top-$k$ certification and simultaneous rank intervals. Extensive experiments on synthetic data and four real-world benchmarks demonstrate that BB-EDGE maintains anytime FWER control while achieving high efficiency.