AI 中文总结
针对排行榜无法证明系统可比性的问题,提出组成可控性检验方法,并构建BioLitBench基准,证明SCRIBE在匹配证据下获得认证排名优势。
AI 中文摘要
排行榜根据长时程智能体的最终输出对其进行排名。然而,仅凭更高的分数并不能确定两个系统是否具有可比性,也无法确定哪个阶段造成了差异。不平等的证据、输入或预算会影响分数,而统计校正并不能消除这种不匹配。我们引入组成可控性来解决这些问题。比较窗口覆盖一个阶段、多个阶段或整个智能体。我们的核心结果仅利用窗口外的干扰因素来界定观测到的分数差异与受控分数差异之间的差距。这产生了一个在检查分数之前应用的容许性检验。不容许的比较会被拒绝。对于容许的比较对,只有当分数差距超过采样和干扰半径之和时,排序才被证明;否则,该排序仍处于未决状态。这些决策为每个系统提供了一个排名区间。我们引入了BioLitBench,一个包含2,042篇生物医学文章的结构化声明图基准。在七个已发表的流程中,传统的统计分析在21对两两比较中宣布了14个获胜者。然而,排名最高的系统独自获得了目标综述的参考文献列表。为了隔离流程性能,我们的检验要求匹配的输入和固定的骨干模型。它拒绝了21个比较中的11个,包括所有涉及排名最高系统的比较。14个传统结论中有7个落在这些被拒绝的比较对中。相同的比较窗口支持阶段级训练。我们使用在每个阶段出口处测量的奖励在Qwen3.8-27B上训练SCRIBE。在匹配的证据下,SCRIBE获得了[1,2]的认证排名区间,并相对于所有评估的已发表流程以及评估的Claude和OpenAI智能体具有认证优势。在相同池下,SCRIBE匹配最强的已发表检索器,并被认证高于三个已发表流程。
英文摘要
Leaderboards rank long-horizon agents by their final outputs. Yet a higher score alone does not establish whether two systems are comparable or which stage accounts for the difference. Unequal evidence, inputs, or budgets can affect scores, and statistical corrections do not remove this mismatch. We introduce compositional controllability to address these questions. A comparison window covers one stage, several stages, or the whole agent. Our central result bounds the gap between observed and controlled score differences using only nuisance outside the window. This yields an admissibility test applied before scores are inspected. Inadmissible comparisons are refused. For admissible pairs, an ordering is certified only when the score gap exceeds the combined sampling and nuisance radii; otherwise, it remains undecided. These decisions give each system a rank interval. We introduce BioLitBench, a benchmark of 2,042 biomedical articles represented as structured claim graphs. Among seven published pipelines, a conventional statistical analysis declares a winner in 14 of 21 pairwise comparisons. Yet the top-ranked system alone received the target review's bibliography. To isolate pipeline performance, our test requires matched inputs and a fixed backbone model. It refuses 11 of the 21 comparisons, including every comparison involving the top-ranked system. Seven of the 14 conventional conclusions fall within these refused pairs. The same comparison windows support stage-level training. We train SCRIBE on Qwen3.8-27B using rewards measured at each stage's exit. Under matched evidence, SCRIBE achieves a certified rank interval of [1,2], with certified advantages over all evaluated published pipelines and the evaluated Claude and OpenAI agents. Under same pool, SCRIBE matches the strongest published retriever and is certified above three published pipelines.
Comments33 pages, 1 figure, 18 tables. Zhaowei Han, Xiang Zhang, and Lingxiao Guan contributed equally. Code: https://github.com/shawnzhg/LitReview-SCRIBE