发表机构
Novelis Research(诺贝丽斯研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究量化小语言模型推理时轨迹问题,提出MGT - B方法,通过映射窗口到概率、累积因子等操作监测并响应警报,经实验得出该方法对特定设置有一定效果,支持选择性监测与修复机制。
AI 中文摘要
量化的小自回归推理模型可能会进入冗长、重复或无成效的轨迹,而推理时的计算分配通常未考虑轨迹的发展情况。基于早期的令牌级电子累积和控制器,我们开发了MGT-B(监测引导测试时回溯),这是一种经过修订的外部控制器,它将预采样不确定性和退化特征的重叠窗口映射到位置条件经验尾部概率,通过CUSUM形状的重置累积混合投注因子,并通过估计回滚点、恢复令牌和键值缓存状态以及执行受限重解码来响应警报。为了审核在手动选择日志阈值h = 10后首次观察到的问题标识上效果是否持续,我们回顾性地排除了阈值前工件中存在的260个标识,并为每个剩余标识保留按时间顺序排列的第一个阈值后对,从而产生一个240对的时间顺序审核集。在这个集合上,准确率从82/240变为88/240(提高2.50个百分点;13次校正,七次回归;精确的麦克尼马尔检验p = 0.2632;配对自举95%区间[-1.25,+6.25])。一个更广泛的467对种子匹配对的历史覆盖集将准确率从146/467变为167/467(提高4.50个百分点;麦克尼马尔检验p = 0.000753),但包括200个在阈值选择之前或期间可用的种子1标识,并且仅作为探索性估计报告。467对集合中的所有316个无警报输出与普通输出相同,而151个有警报的轨迹包含29次校正和8次回归。两种分析都不是确定性的,并且经验因素未被确立为有效的电子过程或电子探测器。结果支持针对所研究的MATH - 500设置的选择性监测和修复机制,而不是一般性的或理论上经过认证的推理改进。
英文摘要
Quantized small reasoning models can enter repetitive or otherwise unproductive trajectories, yet standard decoding does not adapt to the trajectory as it unfolds. We study MGT-B, a fixed, weight-preserving controller that converts overlapping windows of uncertainty, repetition, and local-change features into position-conditional empirical tail probabilities. It accumulates mixture betting factors with a CUSUM-shaped reset, and, after an alarm, restores a coherent earlier token and key-value-cache state before constrained re-decoding. On MATH-500, a paired three-seed evaluation over 1,500 generations per method raises exact-normalized accuracy from 54.73% for vanilla decoding to 56.40% (+1.67 percentage points; problem-clustered bootstrap 95% CI [+0.47, +2.80]), while a prospectively profiled random-intervention control reaches 54.60%. The gain is positive in all three seeds and costs 5.14% more sampled tokens. Seed-0 ablations show that rollback alone does not explain the result and that an isolated repetition penalty is harmful. Five-sample self-consistency reaches 70.0% but uses about 4.84x as many tokens as MGT-B. On the harder, non-overlapping Omni-MATH evaluation, however, MGT-B obtains 16.60% versus 16.67% for vanilla (-0.07 points; clustered 95% CI [-0.33, +0.20]) with 2.10% more sampled tokens. Thus, MGT-B provides a modest, reproducible local improvement on MATH-500 in the studied configuration, but the effect does not transfer to Omni-MATH and should not be interpreted as a general improvement in mathematical reasoning.