发表机构
College of Aeronautics and Engineering, Kent State University(肯特州立大学航空与工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过输入盲对照实验,证明多项选择评估中层程序的巨大神谕余量可能并非源于所选计算的特有收益,提示需谨慎解读自适应层选择的优势。
AI 中文摘要
自适应计算旨在通过针对每个输入定制执行来改进语言模型推理。对于层程序,神谕评估利用已知答案来估计这种灵活性带来的潜在收益,此时实际选择器尚不可用。然而,选择带来的收益本身并不能解释所选程序为何有效。本研究使用两个模型上的32个跳层和重复程序以及4,413个多项选择项目来考察这一区别。分析将这些程序相对于在无评估提示情况下选择的固定动作的收益,与在同一位置进行输入盲扰动(在另一提示下重新评估选择)的收益进行了比较。在共享选项顺序下,对照在Qwen3-4B-Base和Llama-3.1-8B上分别产生10.2-11.8和15.6-19.4个百分点的余量,在每个模型的全部三次随机方向抽取中均超过真实程序的9.0和10.1个百分点。它们仅匹配答案变化率,且排序取决于菜单:在事后比较中,真实程序在Llama的仅重复菜单上每次抽取均领先。一个较小的KL校准比较(包括一个输入相关对照)在点估计上有利于真实程序,但校正检验不具结论性。固定字母偏移产生相似规模的余量。旋转选项大幅降低两类程序的余量,同时留下1.4-2.3和3.7-4.5个百分点的正的真实减对照差异;其幅度和统计支持取决于进一步调整和参考。一项补充的生成答案测试发现,搜索选择的程序在改写后仍保持对为其他问题选择的程序26.0个百分点的优势,但无安慰剂对照。这些结果表明,在共享选项顺序下,大量余量可跨提示持续存在,但并未确立所选层计算特有的收益;与这些对照的排序均无法识别该收益。
英文摘要
Adaptive computation aims to improve language-model inference by tailoring execution to each input. For layer programs, oracle evaluations use known answers to estimate the potential gain from this flexibility, before a practical selector is available. However, a gain from selection does not by itself explain why the chosen programs help. This study examines this distinction using 32 layer-skipping and repetition programs on two models and 4,413 multiple-choice items. The analysis compares their gains over a fixed action selected without the evaluation prompt with those of input-blind perturbations at the same sites, re-evaluating selections on another prompt. With shared option order, the controls give 10.2-11.8 and 15.6-19.4 percentage points of headroom on Qwen3-4B-Base and Llama-3.1-8B, exceeding the real programs' 9.0 and 10.1 in all three random-direction draws per model. They match answer-change rate only, and the ordering depends on the menu: in post hoc comparisons, real programs lead on Llama's repeat-only menu in every draw. A smaller KL-calibrated comparison, including an input-dependent control, favours real programs in point estimate, with inconclusive corrected tests. Fixed letter offsets produce headroom of similar scale. Rotating options sharply reduces both families' headroom, while leaving positive real-minus-control differences of 1.4-2.3 and 3.7-4.5 points; their magnitudes and statistical support depend on further adjustments and the reference. A supplementary generated-answer test finds that search-selected programs keep a 26.0-point advantage over programs selected for other problems after rewording, without a placebo comparison. These results show that substantial headroom can persist across prompts with shared option order without establishing a benefit specific to the selected layer computation; neither ordering against these controls identifies that benefit.