arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

哪种模型表现最佳取决于你给的时间:大语言模型评估中依赖预算的排名

Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation

Rodrigo Guedes de Souza, Alison R. Panisson

arXiv 2608.12150首次发表:更新:

发表机构

Federal University of Santa Catarina (UFSC)(圣卡塔琳娜联邦大学(UFSC))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过调整 token 生成预算评估 4 个大语言模型在 3 个推理基准的表现,发现模型排名随预算变化反转,提出需采用预算条件化的评估协议。

AI 中文摘要

大语言模型的标准评估假设在不同推理条件下模型排名稳定。我们通过调整 token 生成预算(即模型可生成的最大 token 数),设置 7 个水平(64 至 4096),在 3 个推理基准上评估 4 个模型,共完成 56476 次推理,以挑战该假设。我们报告四项发现:(i)即使控制截断,3%-19% 的条目呈现非单调行为(准确率随预算增加而下降),该现象具有模型特异性(跨模型重叠度为 6%-14%);(ii)所有基准上模型排名随预算变化出现反转(p<0.01,采用 McNemar 检验);(iii)Oracle 分析显示模型互补性最高可达 27.8 个百分点,在受限预算下最为显著;(iv)感知预算的路由器可跨域捕捉 14.1% 的 Oracle 差距,预算特征在域内有帮助(提升 1.6 至 5.7 个百分点)但具有域特异性,会损害迁移性能(下降 1.2 个百分点)。这些结果主张采用考虑预算的评估协议。

英文摘要

Standard evaluation of large language models assumes stable model rankings across inference conditions. We challenge this assumption by varying the token generation budget, i.e., the maximum tokens a model may produce, across seven levels (64--4,096), evaluating four models on three reasoning benchmarks (56,476 inferences). We report four findings: (i) 3--19% of items exhibit non-monotone behavior (accuracy decreasing with more budget), even after controlling for truncation, and this phenomenon is model-specific (cross-model overlap: 6--14%). (ii) Model rankings reverse across budgets on all benchmarks ($p {<} 0.01$, McNemar). (iii) Oracle analysis reveals model complementarity up to $+27.8$pp, most pronounced at constrained budgets. (iv) A budget-aware router captures 14.1% of the oracle gap cross-domain; budget features help within-domain ($+1.6$ to $+5.7$pp) but are domain-specific and hurt transfer ($-1.2$pp). These results argue for budget-conditioned evaluation protocols.

Comments19 pages, 11 figures, 7 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑