arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SAEVerbalizer:通过表示 verbalization 为稀疏自编码器特征生成解释

SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization

Weihan Meng, Hongzhu Guo, Yi Jing, Dewen Liu, Zijun Yao, Xiaozhi Wang, Lei Hou, Juanzi Li

arXiv 2608.13538首次发表:更新:

发表机构

Tsinghua University; Peking University; Fudan University(清华大学; 北京大学; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出 SAEVerbalizer 框架,通过微调 LLM 下游层生成 SAE 特征的自然语言解释,解决了传统解释方法的表面性与低效率问题,其 verbalization 能力可泛化、迁移并扩展到不同 LLM 的 SAE 特征。

AI 中文摘要

稀疏自编码器(SAE)被提出用于从大语言模型(LLM)的表示中提取大量特征,但解释这些特征目前仍主要依赖外部观察。这种依赖导致从观察到的模型行为中推断出的解释较为表面,以及大规模收集此类行为证据时的计算效率低下。我们提出 SAEVerbalizer,这一框架将 SAE 解码器方向注入 LLM 的表示中,并微调 LLM 的下游层以生成所注入特征的自然语言解释。训练完成后,所得 verbalizer 可直接从解码器方向解释 SAE 特征,解决上述两个局限。我们的实验表明,所学 verbalization 能力可泛化到未见过的特征,在分别训练的 SAE 词典间迁移,且借助轻量级适配器可扩展到来自不同 LLM 的 SAE 特征。干预实验显示,注入多个方向会产生结合其含义的解释,而反转单个方向则会产生相应的含义转变。

英文摘要

Sparse autoencoders (SAEs) are proposed to extract numerous features from large language model (LLM) representations, yet explaining these features still relies primarily on external observation. This reliance leads to superficial explanations inferred from observed model behavior and computational inefficiency from collecting such behavioral evidence at scale. We introduce SAEVerbalizer, a framework that injects SAE decoder directions into an LLM's representations and fine-tunes the LLM's downstream layers to generate natural-language explanations of the injected features. Once trained, the resulting verbalizer explains SAE features directly from decoder directions, addressing both limitations. Our experiments show that the learned verbalization capability generalizes to unseen features, transfers across separately trained SAE dictionaries, and, with a lightweight adapter, extends to SAE features from different LLMs. Intervention experiments show that injecting multiple directions yields an explanation combining their meanings, while reversing individual directions produces corresponding meaning shifts.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑