超越深度截断:递归语言模型中深度利用的受控评估
Beyond Depth Truncation: Controlled Evaluation of Depth Utilization in Recursive Language Models
浏览论文内容
中文总结 AI 辅助
针对递归语言模型深度利用评估中深度截断指标混淆多种因素的问题,提出深度控制协议(DCP),通过阳性、阴性及训练对照解耦各因素,验证深度利用的真实性。
中文摘要 AI 辅助
深度递归语言模型迭代地应用一个小的层堆栈,将每个词元的计算量与不同的参数数量解耦。为了确定这样的模型是否真正利用其深度,递归和层剪枝文献都依赖于一种共同的评估方法:在推理时截断深度,绘制质量与保留深度比例的关系图,并读取斜率。虽然这种方法廉价且无需训练,但它存在一个未经审视的缺陷:它从同时改变模型多个属性的干预中提取单个标量。深度截断同时减少了块应用的数量,减少了所执行的不同计算量,并将读出头推入分布外的残差流。观察到的斜率混淆了这三个因素,但通常被解释为仅反映第二个因素。我们提出了深度控制协议(DCP),这是一个诊断套件,用于解开这三个量。DCP包括三个阳性对照,在变化其他因素的同时隔离每个因素,一个阴性对照,将相同的干预应用于密集Transformer以确保效果不是测量协议的伪影,以及一个受控的训练干预以验证因果关系。关键对照,即运行完整的块应用预算而仅执行单个不同迭代,仅在深度共享权重的架构中严格可实现,因为在密集网络中重复一层会产生完全不同的模型,而不是同一模型在替代配置中的形式。
英文摘要
Depth-recurrent language models iteratively apply a small layer stack, decoupling per-token compute from distinct parameter count. To determine whether such a model genuinely utilizes its depth, both recurrence and layer-pruning literatures rely on a shared evaluation: truncating depth at inference time, plotting quality against retained depth fraction, and reading off the slope. While cheap and training-free, this metric suffers from an unexamined flaw: it extracts a single scalar from an intervention that alters multiple model properties simultaneously. Depth truncation concurrently reduces the number of block applications, decreases the volume of distinct computation performed, and pushes the readout head onto an out-of-distribution residual stream. The observed slope conflates all three factors, yet is conventionally interpreted as reflecting solely the second. We propose the Depth Control Protocol (DCP), a diagnostic suite that disentangles these three quantities. DCP comprises three positive controls that isolate each factor while varying the others, a negative control applying the identical interventions to dense transformers to ensure the effect is not an artifact of the measurement protocol, and a controlled training intervention to verify causality. The linchpin control, running the full budget of block applications while executing only a single distinct iteration, is strictly realizable only in depth-wise weight-sharing architectures, since in a dense network repeating a layer yields an entirely different model rather than the same model in an alternative configuration.
发表机构
- Blaze AI
- NuverxAI - AI & Creative Innovation Company Limited(NuverxAI - 人工智能与创意创新有限公司)
- Ho Chi Minh City University of Technology (HCMUT)(胡志明市理工大学)
机构由 AI 辅助整理,请以论文原文为准。