arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

仅预填充决策模型中的读出稳定性:零标签预测与推理时计算分配

Readout Stability in Prefill-Only Decision Models:Zero-Label Prediction and Inference-Time Compute Allocation

Ran Li, Lei Chen

arXiv 2610.07716首次发表:更新:

发表机构

Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文证明仅预填充决策模型具有读出稳定性,可利用首遍分布零标签预测菜单干预效果,从而在部署前优化推理时计算分配。

AI 中文摘要

受Jev模型启发的仅预填充决策模型在单次前向传播中对菜单中的每个候选进行评分,且从不解码,这使得一次调用比同规模生成式语言模型便宜一到两个数量级。我们证明这种读出结构具有一个可测试的性质。当干预仅改变候选菜单而保持输入文本不变时,干预后的准确率已由缓存的首遍分布决定。该估计器将首遍概率限制在菜单上,重新归一化,并读出argmax;它不使用标签,也不进行第二次前向传播。在七个模型家族、十个数据集和两种任务类型上,仅菜单干预的预测误差在4.2个百分点以内,且对于一个家族,预测是精确的。同一估计器的概率级变体误差为21.0个百分点,因此该性质存在于排序而非概率中,且无法通过校准恢复。同规模生成式语言模型不具备该性质。在这些模型上,同一估计器的误差为1.6至15.8个百分点,并随模型增大而恶化。该性质将推理时计算转化为可在部署前做出的决策。均匀的额外传递能带来校准但几乎不提高准确率;在匹配成本下,置信级联优于任何重新询问同一模型的方案,且策划菜单优于扩大模型,一个0.8B模型在策划的5候选菜单上达到95.4%的CLINC150准确率,而4B模型在完整150标签上仅为80.0%。代码和数据可在https://this https URL获取。

英文摘要

Prefill-only decision models inspired by the Jev model score every candidate in a menu during a single forward pass and never decode, which makes one call one to two orders of magnitude cheaper than a same-scale generative language model. We show that this read-out structure comes with a testable property. When an intervention changes only the candidate menu and leaves the input text fixed, the post-intervention accuracy is already determined by the cached first-pass distribution. The estimator restricts the pass-1 probabilities to the menu, renormalizes, and reads off the argmax; it uses no labels and no second forward pass. Across seven model families, ten datasets and two task types, menu-only interventions are predicted to within 4.2 points, and for one family the prediction is exact. A probability-level variant of the same estimator errs by 21.0 points, so the property lives in the ranking rather than in the probabilities and is not recovered by calibration. Same-scale generative language models do not share the property. On those models the same estimator errs by 1.6 to 15.8 points and degrades as the model grows. The property turns inference-time compute into a decision that can be made before deployment. Uniform extra passes buy calibration but almost no accuracy; at matched cost a confidence cascade outperforms every scheme that re-asks the same model, and curating the menu beats enlarging the model, with a 0.8B model on a curated 5-candidate menu reaching 95.4% on CLINC150 against 80.0% for a 4B model on the full 150-label menu.Code and data are available at https://github.com/rlisml/jev-cascade.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑