AI 中文总结
该研究构建奖励感知稀疏自编码器RI-SAE,发现其对Llama-3.1-8B的好坏推理延续分离主要源于解决方案完备性,仅少数特征与推理相关,为利用RL信号做模型可解释性提供参考。
AI 中文摘要
稀疏自编码器(SAE)将语言模型的激活分解为稀疏且可解释的特征,一种将其用于推理任务的诱人方法是利用强化学习已产生的奖励信号来整理数据。我们构建了此类奖励感知SAE(RI-SAE):将GRPO轨迹划分为高奖励(“好”)和低奖励(“差”)的推理延续部分,在其激活上训练标准JumpReLU SAE,随后探究所得的好坏分离实际测量的是什么。在Llama-3.1-8B上,16384个特征中的稀疏子集确实能分离两类(所选特征的轮廓系数为0.79,而完整代码的轮廓系数为0.005),但对照实验显示,这种分离主要是解决方案完备性而非推理质量:TF-IDF文本分类器已能分离两类(AUC为0.75至0.83),仅三个结构线索(长度、封闭推理块、带框答案)就达到AUC 0.70(99%的好延续和69%的差延续为带框答案)。从未见过奖励的通用SAE完全无法分离两类(轮廓系数为0.01,无判别特征),因此0.79是对整理信号的样本内拟合,而非无奖励词典恢复的结构。我们因此将该方法及其对照实验一并呈现:奖励过滤是一种廉价、无标签的方式,可复用RL信号用于可解释性,但它揭示的大部分内容是完成形式。仍有两个可解读的判别特征(符号数学;程序性与评价性语言),我们将其视为示例而非孤立的推理。
英文摘要
Sparse autoencoders (SAEs) decompose language-model activations into sparse, interpretable features, and an appealing way to aim them at reasoning is to curate their data with a signal reinforcement learning already produces: the reward. We build such a reward-informed SAE (RI-SAE): we split GRPO trajectories into high-reward ("good") and low-reward ("bad") reasoning continuations, train a standard JumpReLU SAE on their activations, and then ask what the resulting good/bad separation actually measures. On Llama-3.1-8B a sparse subset of the 16,384 features does separate the classes (silhouette 0.79 on the selected features versus 0.005 for the full code), but a control battery shows the separation is largely solution completeness rather than reasoning quality: a TF-IDF text classifier already splits the classes (AUC 0.75--0.83), and three structural cues alone (length, a closed reasoning block, and a boxed answer) reach AUC 0.70 (99% of good versus 69% of bad completions are boxed). A generic SAE that never saw the reward does not separate the classes at all (silhouette 0.01, no discriminative features), so the 0.79 is in-sample fitting of this curated signal rather than structure that a reward-blind dictionary recovers. We therefore present the recipe and its control battery together: reward filtering is a cheap, label-free way to reuse RL signals for interpretability, but most of what it surfaces is completion form. Two discriminative features are still readable (symbolic mathematics; procedural and evaluative language), which we take as illustrative rather than as isolated reasoning.
Comments6 pages, NEMI