arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

稀疏读出棱镜:在特征而非 token 层面解释 Logit-Lens 分数

Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens

Matteo He, William F. Shen, Xinchi Qiu, Nicholas D. Lane

arXiv 2609.01936首次发表:更新:

发表机构

University of Cambridge(剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出稀疏读出棱镜(SRP),用于独立于拟合语料库分析语言模型读出结构,其稀疏近似重构 logit 差值的能力优于基线,且主导读出特征不受拟合语料库影响,为透镜分析提供了可靠对照。

AI 中文摘要

语言模型对下一个 token 的预测会在各层中逐步形成,透镜(lens)方法通过将中间隐藏状态解码为 token 来追踪这一过程。但透镜读取结果同时反映了隐藏状态与用于解码的读出(unembedding 矩阵)。许多透镜在语料库上拟合,我们发现仅拟合语料库不同的两个透镜对相同隐藏状态会报告不同的 token,我们将这种依赖性称为语料库条件性。为了独立于拟合语料库研究读出结构,我们提出稀疏读出棱镜(Sparse Readout Prism,SRP),它仅利用读出的权重分解读出结构,并将任意 token 的 logit 或 logit 差值表示为稀疏读出特征贡献的总和。这将读出特征揭示为透镜读取分析的新单元,可揭示 token 身份会掩盖的结构,支持在 token、上下文、层及透镜间进行比较。用 SRP 的稀疏近似替换原始读出,可重构出比基于读出行几何关系的 6 个基线中最强者多 8.9-17.3 个百分点的测试 logit 差值;对特征进行消融会使 logit 差值按其 SRP 贡献比例变化。尽管 token 读取结果随拟合语料库变化,但主导读出特征保持稳定。由于 SRP 的构建不使用语料库,它为透镜分析提供了独立于拟合语料库的对照。

英文摘要

A language model's prediction of its next token develops across layers, and lens methods track this process by decoding intermediate hidden states into tokens. But a lens reading reflects both the hidden state and the readout (the unembedding matrix) used to decode it. Many lenses are fit on a corpus, and we show that two lenses differing only in their fitting corpus can report different tokens for the same hidden states. We call this dependence corpus conditionality. To examine readout structure independently of the fitting corpus, we introduce Sparse Readout Prism (SRP), which decomposes the readout using only its weights and expresses any token logit or logit difference as a sum of contributions from sparse readout features. This reveals readout features as a new unit of analysis for lens readings, exposing structure that token identities can obscure and enabling comparisons across tokens, contexts, layers, and lenses. Replacing the original readout with SRP's sparse approximation reconstructs 8.9-17.3 percentage points more of the tested logit differences than the strongest of six baselines built on geometric relations among readout rows. Ablating features shifts logit differences in proportion to their SRP contributions. Although token readings vary with the fitting corpus, the dominant readout feature remains stable. Because SRP uses no corpus in its construction, it provides a control independent of the fitting corpus for lens analyses.

Comments55 pages, 33 figures, 42 tables. Under review. Code: https://github.com/hematteo/sparse-readout-prism Dictionaries: https://huggingface.co/hematteo/sparse-readout-prism

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑