arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大规模语言模型隐藏状态的解释

Interpreting Language Model Hidden States at Scale

Jordan Pettyjohn, Mansi Sakarvadia, Nathaniel Hudson, Daniel McKenzie, Kyle Chard, Ian Foster

arXiv 2608.10260首次发表:更新:

AI 中文总结

研究人员提出OmniLens,一种可适配任意模型宽度激活的Lens方法,通过低秩翻译器和Subset-KL技术降低成本,在LLaMA-3.3-70B上构建482个Lens集成,以更低成本复现提示注入检测等案例研究的关键结果。

AI 中文摘要

Lens方法通过将大型语言模型(LLM)的中间激活映射到输出词汇表,来解释LLM,揭示下一个token预测如何在网络中形成。已训练的Lens仍然成本高昂:仿射翻译器的参数随模型宽度呈二次增长,而精确的全词汇表Kullback-Leibler(KL)训练则主导内存消耗。因此,此前已训练的Lens仅被应用于参数最多200亿的模型,且仍与特定组件类型绑定。我们提出OmniLens,它将单一Lens系列应用于任意模型宽度的激活,无论是残差流、注意力还是MLP,并结合两种独立的缩放技术。第一,低秩翻译器使每个Lens的参数增长与模型宽度呈线性关系,可减少多达98.4%的可训练参数。第二,Subset-KL仅实现选定的词汇表logits:其Top-k模式可将峰值训练内存减少多达70%,而其重要性采样变体则为完整KL保留无偏随机梯度。这些节省使得能够为LLaMA-3.3-70B构建482个Lens的密集集成,在相同深度下提供残差流设计6倍的覆盖范围。随后,全模型覆盖揭示了单一组件Lens无法揭示的内容:某一行为最明显的组件未必是干预最有效的组件,且最有效的干预位于此前Lens研究未考察的注意力头之外。在三个案例研究(提示注入检测、多跳记忆注入和毒性定位)中,OmniLens以显著更低的成本复现了已发表的关键结果。

英文摘要

Lens methods interpret large language models (LLMs) by mapping intermediate activations to the output vocabulary, revealing how next-token predictions develop through the network. Trained lenses remain expensive: affine-translator parameters grow quadratically with model width, while exact, full-vocabulary Kullback--Leibler (KL) training dominates memory. Consequently, prior trained lenses have been applied to models of at most 20B parameters and remain tied to particular component types. We present OmniLens, which applies a single lens family to any model-width activation, whether residual stream, attention, or MLP, and combines two independent scaling techniques. First, low-rank translators make per-lens parameter growth linear in model width and reduce trainable parameters by up to 98.4%. Second, Subset-KL materializes only selected vocabulary logits: its Top-k mode cuts peak training memory by up to 70%, while its importance-sampled variant retains unbiased stochastic gradients for the full KL. These savings enable a dense ensemble of 482 lenses for LLaMA-3.3-70B, providing 6x the coverage of a residual-stream design at the same depth. Model-wide coverage then reveals what single-component lenses cannot: the components where a behavior is most visible need not be those where intervention is most effective, and the most effective interventions lie outside the attention heads examined by prior lens studies. Across three case studies (prompt-injection detection, multi-hop memory injection, and toxicity localization), OmniLens reproduces key published results at substantially lower cost.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑