具有更丰富反馈的基准测试的可复用性如何?
How Reusable Are Benchmarks with Richer Feedback?
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究探讨丰富反馈下基准测试的可复用性,发现多标准自适应选择模型时所需测试集大小随标准数指数增长,攻击多任务大语言模型基准可产生虚假胜出者,挑战了现有可靠复用解释。
AI中文摘要:
我们研究当开发者在多个标准下适应评估反馈时,基准测试能否可靠地指导模型选择。我们发现,在任意标准凸组合下,估计从$k$个自适应选择的模型中最佳得分所需的最坏情况测试集大小,随标准数量呈指数增长,在固定准确度和置信度下,仅需$O(\log k)$个标准即可达到回答$k$个自适应统计查询的$\Theta(\sqrt{k})$成本。在对具有五到十个标准的多任务大语言模型基准测试的攻击中,仅限于非支配任务配置文件的反馈会产生较大的复用集与留出集得分差距,并频繁出现虚假的胜出者。这些结果挑战了先前观察到的可靠基准复用现象的一个突出解释——即开发者主要响应于对当前最佳模型的令人信服的改进——在丰富反馈设置中,同时留下了普通模型开发中这种脆弱性出现频率的开放问题。
英文摘要:
We study whether benchmarks reliably guide model selection as developers adapt to evaluation feedback across multiple criteria. We find that the worst-case test-set size needed to estimate the best score among $k$ adaptively chosen models, under any convex combination of the criteria, grows exponentially with the number of criteria, reaching the $Θ(\sqrt{k})$ cost of answering $k$ adaptive statistical queries with only $O(\log k)$ criteria, at fixed accuracy and confidence. In attacks on multi-task large language model benchmarks with five to ten criteria, feedback restricted to nondominated task profiles produces large reused-to-held-out score gaps and frequent false winners. These results challenge a prominent explanation for prior observed reliable benchmark reuse---that developers mainly respond to convincing improvements over the current best---in rich-feedback settings, while leaving open how often ordinary model development encounters this vulnerability.