arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01345cs.AIcs.CRcs.LG

廉价验证器,巨大的盲区:测量成本节约级联的可靠性成本

Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving Cascades

Dushyant Rajput, Nirdesh Chauhan, Siddharth Kosaraju

首次发表
浏览论文内容

中文总结 AI 辅助

该研究发现,基于LLM的成本节约推理级联存在巨大验证器盲区,朴素微调无法改善性能,且级联自身指标无法反映真实错误,得出其可靠性无法通过自身验证器指标判断的结论。

中文摘要 AI 辅助

推理级联通过用廉价模型回答大部分查询,并将困难的尾部查询升级为作为验证器的前沿模型来降低成本。一个自然的扩展是形成闭环:在验证器的弃权(不执行)样本上微调廉价学生模型,从而使每一轮的升级率和成本都下降。我们在真实的大型语言模型(LLM)上测量这个闭环,并报告四项发现。第一,验证器的盲区(即它接受的学生错误答案的比例)很大且具有对抗性:它随学生能力增长而增加(当学生规模从0.5B扩大到32B时,β从0.12升至0.55),随验证器能力增强而缩小,因此在级联存在的廉价学生-廉价验证器机制中,盲区最为严重。第二,消除盲区会损失节约的成本:前沿验证器可将β降至约0.05,但随后会对46%的困难MATH查询进行升级,而此时真实错误率为39%,近一半的流量都需支付前沿模型的成本。第三,对验证器拒绝的尾部进行朴素的校正微调,无法提升小型学生模型,反而会使其性能下降并最终崩溃,无论我们尝试的是跨家族还是同家族的教师模型都是如此,因此在该规模下,这种自我改进的闭环是自我挫败的。第四,在整个过程中,级联自身的仪表盘(所有通过验证器计算的指标)显示错误率稳定在3%,而实际交付的错误率最高可达32%:该系统从设计上就对自身的退化视而不见。随后,我们给出解释这种盲性的理论,即双总体守恒定律ε_∞ ≲ q₀β₀,在该定律下,所有闭环内指标均会改善,而真实质量却不会提升,还通过合成研究验证了该机制。实际结论是:自我改进级联的可靠性无法通过其自身验证器计算的任何指标来判断。

英文摘要

Inference cascades cut cost by answering most queries with a cheap model and escalating a hard tail to a frontier model that acts as verifier. A natural extension closes the loop: fine-tune the cheap student on the verifier's rejections so the escalation rate, and cost, fall each round. We measure this loop on real LLMs and report four findings. First, the verifier's blind spot, the fraction of the student's wrong answers it accepts, is large and moves adversarially: it grows with student capability ($β$ from 0.12 to 0.55 as the student scales 0.5B to 32B) and shrinks with verifier capability, so it is worst in the cheap-student, cheap-verifier regime cascades exist to create. Second, buying it away returns the saving: a frontier verifier drives $β$ to about 0.05 but then escalates on 46% of hard-MATH queries against a 39% true error rate, paying the frontier price on nearly half of all traffic. Third, naive corrective fine-tuning on the verifier-rejected tail does not improve the small student but degrades and ultimately collapses it, across every teacher we tried (cross-family and same-family), so at this scale the self-improving loop is self-defeating. Fourth, through all of this the cascade's own dashboard, every metric computed through the verifier, reads a flat 3% error while true delivered error swings up to 32%: the system is blind to its own degradation by construction. We then give the theory that explains the blindness, a two-population conservation law, $ε_\infty \lesssim q_0 β_0$, under which every in-loop metric improves while true quality does not, and a synthetic study that validates the mechanism. The practical conclusion: the reliability of a self-improving cascade cannot be read from any metric computed through its own verifier.

发表机构

  • AltSlate Labs LLP(奥特斯莱特实验室有限责任公司)

机构由 AI 辅助整理,请以论文原文为准。

↑