arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种校准的测量推理优化对输出质量影响的工具

A Calibrated Instrument for Measuring How Inference Optimizations Affect Output Quality

Jerry Kaplan

arXiv 2609.18005首次发表:更新:

发表机构

Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种校准的LLM评判测量方法,统一评估量化、早退等推理优化对输出质量的影响,发现质量损失因领域和模型而异,确定性不能区分重要错误。

AI 中文摘要

大语言模型优化是一个活跃的研究领域,涵盖模型权重量化、跳过层的早退方法以及投机解码。每条技术路线都使用自己的质量度量,通常是特殊的基准分数。很少有方法能达到其他科学学科所需的测量精度。我们提出了一种严格的输出质量测量方法,适用于跨系统和跨技术的比较。我们使用LLM作为评判者对输出进行评分,但正式地校准评判者:我们比较同一模型在相同提示下两次普通运行的评分,验证其对统计等效输出没有系统性偏好,并测量其每样本噪声。每个设计还包含一个'空'条件,其分布与未修改模型在证明上相同,其测量差异必须为零。使用这一种工具,我们在相同提示下测量了几种加速技术,因此可以比较它们的质量成本。感知质量高度依赖于话语领域。一个4位模型与其16位原始版本在我们的设计±0.3点分辨率内无法区分,在英语散文和中文中都是如此。在3位精度下,相同提示在英语散文中损失0.5点,在中文中损失0.9点,在多步数学上损失1.1点;早退在散文中损失0.7点,在数学上损失2.5点,将正确解决的问题从27个中的19个减少到6个。这种模式在阿里巴巴和Meta的模型中都成立,但幅度不同:相同的量化器使Meta的模型损失1.8点,而阿里巴巴的模型损失0.7点。模型对某个标记的确定性预测了其与完整模型选择不同的可能性,但不能预测这种差异对评判质量的影响程度,因此依赖确定性的接受规则无法区分重要错误和不重要错误。

英文摘要

Large language model optimization is an active research area, spanning quantization of model weights, early-exit methods for skipping layers, and speculative decoding. Each track uses its own quality measures, typically an idiosyncratic benchmark score. Few approach the measurement precision required by other scientific disciplines. We propose a rigorous methodology for measuring output quality, suitable for cross-system and cross-technique comparison. We score outputs with an LLM as a judge, but calibrate the judge formally: we compare its scores on two ordinary runs of a model given the same prompts, verifying that it shows no systematic preference between statistically equivalent outputs and measuring its per-sample noise. Each design also includes a 'null' condition, provably identical in distribution to the unmodified model, whose measured difference must be zero. With this one instrument we measure several acceleration techniques on the same prompts, so their quality costs can be compared. Perceived quality proves highly dependent on the domain of discourse. A 4-bit model was indistinguishable from its 16-bit original down to our design's +/-0.3-point resolution, in English prose and Chinese alike. At 3-bit precision the same prompts lost 0.5 points in English prose, 0.9 in Chinese, and 1.1 on multi-step math; early exit that cost 0.7 points on prose cost 2.5 on math, cutting correctly solved problems from 19 of 27 to 6. The pattern held for models from Alibaba and from Meta, but not its magnitude: the same quantizer cost Meta's model 1.8 points where it cost Alibaba's 0.7. A model's certainty about a token predicts how likely it is to differ from the full model's choice, but not how much that difference affects judged quality, so acceptance rules relying on certainty cannot distinguish errors that matter from errors that don't.

Comments23 Pages. 6 tables in main text,5 tables in appendices. Code, prompts, and result files at https://github.com/jerrykaplan/Calibrated-Instrument

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑