arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33012cs.AI

当配对数量不等于样本量:全配对智能体比较所估计的内容

When Pair Count Is Not the Sample Size: What All-Pairs Agent Comparisons Estimate

Wei-Jung Huang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究揭示智能体基准全配对比较中,配对数量并非有效样本量;在固定排行榜下均值为精确汇总,在随机配置下为二阶U统计量,其不确定性取决于配置数而非配对数,并通过SWE-bench等实验证明需明确固定与抽样成分。

中文摘要 AI 辅助

当一个智能体基准比较排行榜上的每一对条目时,比较的数量可能看起来远大于其背后的独立证据:A对B和A对C都重复使用了A。这种重复使用是否影响推断,取决于分析旨在描述什么。如果排行榜及其结果是固定的,那么全配对均值是这些条目的精确汇总,任何区间必须来自另一个声明的随机性来源。相反,如果条目被视为从未来配置总体中独立同分布抽取的,且配对规则是正则且非退化的,那么相同的均值是一个二阶U统计量,其一阶不确定性取决于配置的数量,而非配对的数量。我们以近乎平局作为运行示例,但这一区分也适用于其他对称配对汇总,只要其正则性条件成立。我们使用一个固定的SWE-bench Verified快照和一个已知真值的精确二元模型来检验这两种解释。在SWE-bench上,考虑共享配置的区间比将配对视为独立的配对独立参考区间宽两倍以上。在精确模型中,当边共享端点时,配对独立的覆盖率远低于名义水平,但对于匹配的独立边则保持在名义水平附近。在另外两个固定排行榜上的结果表明,精确汇总还取决于包含哪些配对以及如何加权。因此,全配对分析必须说明什么是固定的、什么是抽样的,以及如何处理共享条目和配对聚合。

英文摘要

When an agent benchmark compares every pair of leaderboard entries, the number of comparisons can look much larger than the independent evidence behind them: A versus B and A versus C both reuse A. Whether this reuse affects inference depends on what the analysis is meant to describe. If the board and its outcomes are fixed, the all-pairs mean is an exact summary of those entries, and any interval must come from another declared source of randomness. If the entries are instead treated as iid draws from a population of future configurations and the pair rule is regular and nondegenerate, the same mean is an order-two U-statistic whose first-order uncertainty depends on the number of configurations, not the number of pairs. We use near ties as the running example, but the distinction extends to other symmetric pair summaries when their regularity conditions hold. We examine both interpretations using a fixed SWE-bench Verified snapshot and an exact binary model with known truth. On SWE-bench, intervals that accounted for shared configurations were more than twice as wide as a pair-iid reference that treated the pairs as independent. In the exact model, pair-iid coverage fell far below the nominal level when edges shared endpoints but remained near nominal for matched independent edges. Results on two other fixed leaderboards show that exact summaries also depend on which pairs are included and how they are weighted. An all-pairs analysis must therefore state what is fixed, what is sampled, and how it handles shared entries and pair aggregation.

补充信息

↑