arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

测量位置决定测量内容:基于消融的SAE评估中的位置选择

Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation

Valentin Noël

arXiv 2608.13337首次发表:更新:

发表机构

Devoteam(德孚替)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究发现基于消融的SAE评估中,测量位置的选择会影响结果,提出需统一测量位置的评估协议,修正方法仅需一行代码,且随语料库规模扩大问题更严重。

AI 中文摘要

稀疏自编码器(Sparse Autoencoder, SAE)旨在识别语言模型计算的内容,检查隐变量重要性的常用方法是将其关闭并观察变化。但隐变量会在多个token处激活,其效应需在其中一个token处测量。常规做法是在隐变量激活最强的位置测量,该选择几乎从未被报告,且不由实验者决定,而是由被评估的字典决定:更换字典,测量位置会移至不同token。我们表明这并非细节:以Google发布的针对同一模型的两个SAE为例,通过解码器相似度匹配隐变量,即使在两个字典编码几乎相同的隐变量对中,仍有很大比例的隐变量对应不同token。因此,按常规协议比较两个字典时,往往在不同位置进行比较。为分离常规做法与字典的影响,我们从同一初始化训练了6个仅拟合选择不同的自编码器,使每个隐变量含义相同。此类比较中被解读为“这些字典对该隐变量存在分歧”的大部分方差,实际源于位置:当每个字典在同一token处测量时,该方差从7.6%和11.9%降至接近零。更多评估数据无法解决此问题:在16倍语料库规模范围内,字典对测量位置的一致性反而更低,问题随规模扩大而加剧。修正方法仅需一行评估代码。我们给出基于消融的因果数需报告的协议,以确保不同论文间具有可比性,并对5篇已发表论文进行了审计。简言之:未报告测量位置的因果数,既描述了隐变量,也描述了其被测量的token。

英文摘要

Sparse autoencoders are meant to name the things a language model computes, and the usual way to check that a latent matters is to switch it off and see what changes. But a latent fires at many tokens, and the effect has to be measured at one of them. The convention is to measure where the latent fires hardest. That choice is almost never reported, and it is not made by the experimenter: it is made by the dictionary under evaluation. Change the dictionary and the measurement moves to a different token. We show this is not a detail. Take two sparse autoencoders released by Google for the same model and match their latents by decoder similarity: even among the pairs the two dictionaries encode almost identically, they pick different tokens for a large share of them. Two dictionaries compared under the usual protocol are therefore very often compared at different places. To separate the convention from the dictionaries we train six autoencoders from one initialisation, differing only in fitting choices, so that a latent means the same thing in each. Most of the variance such a comparison reads as "these dictionaries disagree about this latent" turns out to be the position instead: it falls from 7.6% and 11.9% of variance to near zero once every dictionary is measured at the same token. More evaluation data does not rescue it. Across a sixteenfold range of corpus sizes the dictionaries agree less about where to measure, not more, so the problem grows with scale. The correction is one line of evaluation code. We give the protocol an ablation-based causal number must report to be comparable across papers, and an audit of five published papers against it. In short: a causal number reported without its position describes the token it was taken at as much as the latent it was taken from.

Comments19 pages, 3 figures. Code and data: https://github.com/vcnoel/sae-artifact

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑