arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于成本校准前沿效用的预算感知大语言模型发现

Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility

Yansen Zhang, Yilu Liu, Tianyu Liu, Jiamin Chen, Xiaokun Zhang, Kai Xie, Qingfu Zhang, Xue Liu, Yiyan Qi, Chen Ma

arXiv 2607.26828首次发表:更新:

发表机构

City University of Hong Kong; Mohamed bin Zayed University of Artificial Intelligence (MBZUAI); International Digital Economy Academy (IDEA)(香港城市大学; 穆罕默德·本·扎耶德人工智能大学; 国际数字经济学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对大语言模型发现的成本问题,提出 CostAda 成本自适应控制器,通过成本校准前沿效用优化预算使用,在多基准测试中以更少预算达到或超越现有模型的质量表现。

AI 中文摘要

大型语言模型正通过对已评估候选对象的推理时搜索,越来越多地支持科学与算法发现。现有自适应发现控制器仅基于分数进展分配信用,尽管提示长度、重试次数和引导调用会导致搜索动作产生不同的 token 成本。我们证明,当前沿(frontier)数量增加且成本差异扩大时,未考虑成本的信用分配会损失几乎所有可获得的质量。在固定的搜索侧 token 预算下,控制器必须判断哪个前沿正在改进,以及其收益是否能证明已产生的成本合理,且要在预算耗尽前做出决策。我们提出 CostAda,一种围绕成本校准前沿效用构建的成本自适应控制器,该效用会相对于已产生的动作成本评估前沿进展,并将信用与剩余预算挂钩。CostAda 利用该信号控制局部探索强度、前沿分配和预算策略干预,使成本与剩余预算成为搜索的驱动因素,而非仅作为核算变量或停止规则。在 16 个基准-主干模型对中,CostAda 能用至多一半的预算达到最强基线的全预算质量;在 GLM-5 和 GPT-5.4 下的全部 8 个基准中,其最终平均质量均为最优。

英文摘要

Large language models increasingly support scientific and algorithmic discovery through inference-time search over evaluated candidates. Existing adaptive discovery controllers assign credit based only on score progress, even though prompt length, retries, and guidance calls cause search actions to incur different token costs. Under a fixed search budget, the controller must decide which frontier to develop and whether its gain justifies the realized cost before the budget is exhausted. We prove that, in the worst case, cost-blind control can forfeit all but a vanishing fraction of attainable quality as frontier count and cost heterogeneity grow. To address this limitation, we introduce \textbf{CostAda}, an adaptive controller built on \emph{cost-calibrated frontier utility}. By valuing progress relative to realized action cost and conditioning credit on the remaining budget, CostAda coordinates local exploration intensity, frontier allocation, and budgeted tactic intervention. Realized action cost and remaining budget therefore shape the search rather than serving only as accounting variables or a stopping rule. Across eight benchmarks with GLM-5 and GPT-5.4, CostAda reaches the strongest baseline's full-budget quality with at most half the budget on thirteen of sixteen benchmark--backbone pairs and achieves the strongest final quality on all sixteen pairs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑