一种尺寸并不适合所有场景:根据部署实际提出的问题设置推理深度
One Size Does Not Fit All: Setting Inference Depth from the Questions a Deployment Actually Asks
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文研究提前退出机制,根据部署的特定问题范围设置推理深度,发现节省因流量类型而异,定制退出阈值有效,且token级保真度指标在可验证领域失效。
AI中文摘要:
Transformer语言模型被训练用于响应任何提示,但每次部署只处理狭窄范围的问题:支持助手看到配送投诉,编码工具看到Python代码。然而,每次部署在每token上仍支付相同的计算量。本文衡量了当提示范围事先已知时,这部分成本中有多少是可以避免的。所研究的机制是提前退出:一个小的、经过训练的组件——称为读出器——被附加到中间层并提出一个token,然后通过置信度测试决定是输出该token还是运行剩余层。模型被冻结,唯一使用的监督是模型在普通流量上的自身输出。报告了三个发现。首先,可实现的节省在很大程度上取决于流量类型:在15亿参数模型的一半深度下,对于算术应用题,96%的token可以提前输出,对于中文解释,这一比例为8%,同时与完整模型在token级保真度上匹配(第三项发现揭示了该指标的局限性)。其次,在部署可能利用其流量知识的三种方式中,只有定制提前退出的阈值是有价值的:针对每次部署校准该阈值,在三个模型上将退出率提高了最多59个百分点,在测试的大多数语料上提高了超过10个百分点。第三,token级保真度——提前退出文献中的标准评估指标——在token可以对照真实值检查的领域中失效:在算术应用题上,三个模型在完整运行时各自正确回答了60个问题,而在提前退出下,在保真度得分最高的配置中,正确回答数在10到28之间。预期场景是个人设备上的小模型,其中生成受限于内存带宽而非计算。
英文摘要:
A transformer language model is trained to respond to any prompt, but each deployment asks only a narrow range of questions: a support assistant sees delivery complaints, a coding tool sees Python. Every deployment nonetheless pays the same computation per token. This paper measures how much of that cost is avoidable when the range of prompts is known in advance. The mechanism examined is early exit: a small, trained component - called a readout - is attached to an intermediate layer and proposes a token, and a confidence test decides whether to emit it or to run the remaining layers. The models are frozen, and the only supervision used is the model's own output on ordinary traffic. Three findings are reported. First, achievable savings depend strongly on the kind of traffic: at half depth on a 1.5-billion-parameter model, 96 percent of tokens could be emitted early for arithmetic word problems and 8 percent for Chinese-language explanations, at matched token-level fidelity to the full model (a measure whose limits the third finding exposes). Second, of three ways a deployment might use knowledge of its traffic, only customizing the threshold for exiting early is worthwhile: calibrating it per deployment raised exit rates by up to 59 percentage points across three models, and by more than 10 points on most corpora tested. Third, token-level fidelity - the standard evaluation measure in the early-exit literature - fails in domains where tokens can be checked against ground truth: on arithmetic word problems, three models each answered sixty questions correctly when run in full, and between 10 and 28 correctly under early exit, in the configuration that scored highest on fidelity. The intended setting is small models on personal devices, where generation is limited by memory bandwidth rather than computation.