arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

哪种大语言模型适用于哪种工作?不确定评估下的预算模型分配

Which LLM for Which Work? Budgeted Model Allocation under Uncertain Evaluation

Hamed Khosravi, Xiaoming Huo

arXiv 2608.29560首次发表:更新:

发表机构

Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对固定AI预算下LLM分配的不确定评估问题,提出CASE方法,通过双求解证书判断最优分配,实验表明模型质量信息优化比分配优化更具价值。

AI 中文摘要

一家拥有固定人工智能(AI)预算的公司必须决定由哪个大语言模型(LLM)处理每一项 recurring workload( recurring workload 可译为“ recurring 工作负载”,此处保留英文原名)。该公司所欠缺的是质量表,即每个模型在每项工作负载上的表现如何。若有了该表,此决策就是一个多选背包问题,求解是常规操作,因此估计该表才是难点,而这种估计存在两个问题。模型很少在相同工作上进行比较,且记录的分数通常是代理指标,而非公司看重的结果。因果和离策略方法可解决第一个问题,但以第二个问题为条件;而评估者验证方法可估计第二个问题,但无法完成决策。更糟糕的是,购买更多重新评估无法解决第二个问题:随机化决定了哪些请求被评分,而非分数如何产生,因此无论购买多少评估,该表仍不确定。不过,即使该表不确定,部署决策仍可能确定。因此,我们提出问题:是否存在一种分配方式,在所有与证据一致的质量表下均保持最优?对于固定预算问题,这存在一个精确的双求解证书:在估计表上求解一次,在最不利表上求解一次。一致则证明该分配有效;不一致则识别出模型-工作负载对,其中进一步的证据可能产生影响。我们提出 CASE(因果主动序列实验),其将评估定向到这些对,并随着证据积累重复测试。在生产日志中,测量失败是两个问题中更严重的那个:仅修正分配仍会留下大部分损失,而随机重新评估无法消除它。在我们的实验中,现有证据往往无法确定分配。在付费软件任务上,关于模型质量的更好信息比在相同估计上进一步优化分配能产生更多节省。

英文摘要

A company with a fixed artificial intelligence (AI) budget must decide which large language model (LLM) handles each recurring workload. What it lacks is the quality table, how well each model performs on each workload. Given that table, the decision is a multiple-choice knapsack problem and is routine to solve, so estimating it is the difficulty, and that estimation fails in two ways. Models are rarely compared on the same work, and the recorded score is usually a proxy rather than the outcome the company values. Causal and off-policy methods repair the first but condition on the second, while evaluator-validation methods estimate the second but stop short of the decision. Worse, buying more re-evaluation cannot settle the second: randomization governs which requests are scored, not how a score is produced, so the table stays uncertain however much evaluation is purchased. Yet the deployment decision may still be determined even when the table is not. We therefore ask whether one assignment stays optimal across every quality table consistent with the evidence. For the fixed-budget problem, this admits an exact two-solve certificate: solve once at the estimated table and once at a least-favourable table. Agreement certifies the assignment; disagreement identifies the model-workload pairs where further evidence can matter. We propose CASE (causal active sequential experimentation), which targets evaluation to those pairs and repeats the test as evidence accumulates. On a production log, the measurement failure is the larger of the two: correcting assignment exactly still leaves most of the loss, and randomized re-evaluation does not remove it. In our experiments, the available evidence often does not determine the assignment. On paid software tasks, better information about model quality yields more savings than further optimization of the assignment on the same estimates.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑