发表机构
University of Washington; School of Computer Science and Engineering(华盛顿大学; 计算机科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究如何让Transformer清晰可读,提出用通道方差下限等方法,构建出最清晰可读的Transformer,使单元能分离检测与命名,编辑更局部,还揭示清晰度调节旋钮,质量与传统基线相当。
AI 中文摘要
可以通过构造上清晰可读的算子构建Transformer,这些算子是有界的、命名的单元,可视为模糊集操作而非密集激活。但在训练过程中必须强调清晰度,且这种强调存在失败模式。一个清晰度惩罚会将有界算子锐化为决定性检测器,却使其坍缩为死常量。一个等式表明了原因,进而提出修复方法:每个通道的方差下限,作为损失的目标清晰度度量,可恢复清晰度和质量。一个学习到的单位分数取代了先前工作中的手动保留GELU划分。结果是构建出了最清晰可读的Transformer,其前馈操作数的78%和注意力值通道的50%是清晰且上下文相关的检测器,每层的清晰度从浅层的18%提升到深层的78%。这些单元能将清晰检测与更难的命名分离,编辑更具局部性。最后,单元间的去相关压力揭示了清晰度调节旋钮,在不损失质量的情况下,将概念转化为单个可编辑单元,使预测成为简短解释。质量与传统基线相当。
英文摘要
A transformer can be built from operators that are legible by construction -- bounded, named units that read as fuzzy set operations rather than dense activations -- but legibility must be pressed for during training, and the pressure has a failure mode. A crispness penalty meant to sharpen a bounded operator into a decisive detector instead collapses it into a dead constant. An identity, E[v(1-v)] = mu(1-mu) - var, shows why -- the penalty is a variance-minimizer blind to the difference between a live detector and a constant -- and names the fix: a per-channel variance floor, the target legibility metric written as a loss, which recovers both legibility and quality. A learned per-unit fraction then retires the hand-set reserved-GELU partition of prior work: given the choice the model keeps no unit as pure GELU and routes 87% of its load-bearing computation through crisp operators. The result is the most legible transformer we have built -- 78% of its feed-forward operands and 50% of its attention value channels are crisp-and-contextual detectors, and per-head legibility rises from 18% in shallow layers to 78% in deep ones. Read in the correct rotated per-layer frame, these units separate a clean detection (what a unit responds to) from a harder naming (what its output decodes to); and because the objective makes each unit crisp and sparse, edits to them are far more local -- 50-184x in the deep layers where the edit sites concentrate -- and can target explicit conjunctions a single neuron cannot express. Finally, a between-unit decorrelation pressure exposes a legibility dial: it trades a circuit's reuse for independence at no quality cost, turning concepts into single, surgically editable units and a prediction into a short explanation read off a handful of named operations. Quality holds at parity with a conventional baseline throughout.