AI 中文总结
研究在固定问题集上顺序评估新大语言模型,基于历史数据构建置信序列,提出增长导向查询规则,分析影响收缩率因素并提出混合查询规则,实验发现简单的均匀采样有时表现更好。
AI 中文摘要
我们研究了使用先前大语言模型(LLM)的历史性能数据在固定问题集上顺序评估新LLM的问题。目标是构建模型在此问题集上能力的置信序列(CS)并设计能尽快缩小CS宽度的主动查询规则。构建CS时,我们反转一族测试上鞅,关注基于反向信息投影(RIPr)和基于投注测试的方法。先在神谕设置下研究这些方法,证明基于RIPr构建的神谕最优性。接着提出增长导向查询规则,在实际中基于历史数据学习的问题级正确性预测构建测试上鞅和查询规则。分析CS收缩行为,确定减缓收缩率的两个关键因素,最后提出几种混合查询规则减轻这些影响。通过实验比较不同查询规则,有趣的是,简单的均匀采样有时比更自适应的规则表现更好。
英文摘要
We study the problem of sequentially evaluating a new large language model (LLM) on a fixed question set using historical performance data from prior LLMs. Our goal is to construct a confidence sequence (CS) for the model's capability on this question set and to design active querying rules that shrink the CS width as quickly as possible. For CS construction, we invert a family of test supermartingales and focus on two representative approaches: a reverse information projection (RIPr)-based approach and a testing-by-betting-based approach. We first study these approaches under an oracle setting, and demonstrate the oracle optimality of the RIPr-based construction. We then propose a growth-oriented querying rule that aims to maximize the worst-case one-step expected log-increment over the endpoints of the current CS. In practice, we build these test supermartingales and the querying rule on predictions of question-level correctness learned from historical data. We then analyze the shrinkage behavior of the resulting CSs and identify two key factors that slow the shrinkage rate of CSs: accumulated prediction mismatch and the spikiness of the querying distribution. Finally, motivated by this analysis, we propose several mixture querying rules that combine growth-oriented querying, prediction refinement, and uniform exploration, trying to mitigate the effects that slow the shrinkage rate. We provide experiments comparing different querying rules for the RIPr-based and testing-by-betting-based CSs across several synthetic testing datasets. Interestingly, we observe that the simplest querying rule, uniform sampling, can sometimes outperform more adaptive querying rules for both methods.
CommentsThis is a preliminary version; feedback is welcome