arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39277cs.LGcs.CL

倾斜的碗不是滑坡:压缩循环模型

A Tilted Bowl Is Not a Slippery Slope: Compressing Looped Models

Steven Kolawole, Pearse Jim, Opegbemi M. Busoye, Glory Bagai, Virginia Smith

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过实验发现循环模型压缩崩溃源于稳定点偏移而非误差累积,并提出基于暂停头的控制器,在Sudoku-Extreme和Maze-Hard上以更低流量实现更高精度。

中文摘要 AI 辅助

循环模型通过多次应用同一权重块进行推理,因此压缩该块可在每次循环中节省内存流量。然而,压缩后的循环模型常常崩溃,而这种崩溃通常被归咎于循环间累积的舍入误差。在本工作中,我们在五个家族的30多个模型上检验了这一解释,并惊讶地发现,它仅适用于从不稳定的循环。当循环稳定时,固定的舍入误差不会累积。它移动循环稳定的点,就像倾斜碗会移动球静止的位置一样,只有当偏移超过读出容忍度时答案才会丢失。这一图景使我们能够通过一次无标签测量预测哪些模型会失败,并解释了失败模型为何能恢复:它们的循环仍然稳定,因此最后几轮8位权重的循环能带回答案。基于这些发现,我们构建了一个控制器,当模型的暂停头触发时停止,然后以8位循环完成。在Sudoku-Extreme和Maze-Hard上,它在不到三分之一的权重流量下比固定深度推理高出最多15个百分点。

英文摘要

Looped models reason by applying the same block of weights many times, so compressing that block saves memory traffic on every loop. Compressed looped models, however, often collapse, and the collapse is usually blamed on rounding error that accumulates from loop to loop. In this work we test that account on more than 30 models from five families and find, to our surprise, that it holds only for loops that never settle. When a loop settles, a fixed rounding error does not accumulate. It moves the point where the loop settles, much as tilting a bowl moves where a ball comes to rest, and the answer is lost only when the shift is larger than the readout tolerates. This picture lets us predict which models fail from a single label-free measurement, and it tells us why failed models recover: their loops still settle, so a few final loops with 8-bit weights bring the answer back. Motivated by these findings, we build a controller that stops when the model's halting head fires and then finishes with 8-bit loops. On Sudoku-Extreme and Maze-Hard it beats fixed-depth inference by up to 15 points under a third of the weight traffic.

发表机构

  • Carnegie Mellon University(卡内基梅隆大学)
  • ML Collective

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑