arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29551cs.AI

HiPACE:稀疏自编码器中特征吸收的分层相界分析与受控评估

HiPACE: Hierarchical Phase-Boundary Analysis and Controlled Evaluation of Feature Absorption in Sparse Autoencoders

  • Hubei University(湖北大学)

机构由 AI 辅助整理,请以论文原文为准。

Jinyuan Zhang, Peng He, Yin Yuan, He Hu, ShengShuo Jiao

AI总结:

针对稀疏自编码器中特征吸收问题,提出闭式相界预测及HiPACE评估协议,验证了父子结构,因果充分性得到确立。

AI中文摘要:

稀疏自编码器(SAEs)将大语言模型(LLM)的激活分解为稀疏字典原子,使得每个不同的概念拥有自己的特征。一种反复出现的行为使这一前提复杂化:特征吸收,即父概念及其子概念——例如水果与{苹果、香蕉、梨}——会坍缩为一个共享的族方向。先前的工作仅凭经验记录了吸收现象;缺失的是对共享方向何时是活跃语义族的成本最优表示的闭式预测。本文填补了这一空白。对于具有$k$个活跃子概念和残差尺度$\alpha$的分层伯努利生成器,$L_0$惩罚重建目标允许一个闭式相界$\lambda_c(k,\alpha)=\alpha^2 k/(k-1)$:高于该相界,纯父概念吸收严格比纯子概念编码更便宜。基于这一相界,我们引入了HiPACE,一种评估协议,用于测试该相界在真实SAE字典中的结构后果——在WordNet族上测量父子解码器结构,在测试未见族之前冻结发现选择的统计量,并将真实族与随机化兄弟零假设进行对比。该相界在其原生机制中证明是尖锐的,在所有30个测试单元上预测合成转变的误差在$\pm15\\%$以内。在Pythia-160m SAE中,父子解码器差距恢复了预测的排序,偏相关高达$-0.93$,该相关性在锁定的保留集上持续存在,并排除了兄弟零假设($p=0.002$)。受控激活组合将理论的活跃子概念计数与恢复的族方向联系起来,残差流干预表明带符号的族方向增加父类别对数几率,在符号翻转时反转,在随机控制下消失——从而在族子空间层面确立了因果充分性。

英文摘要:

Sparse autoencoders (SAEs) decompose LLM activations into sparse dictionary atoms, so that each distinct concept gets its own feature. One recurring behavior complicates this premise: feature absorption, in which a parent concept and its children--fruit and {apple, banana, pear}, say--collapse into a shared family direction. Prior work documents absorption empirically; missing is a closed-form prediction of when the shared direction is the cost-optimal representation of an active semantic family. This paper closes that gap. For a hierarchical Bernoulli generator with $k$ active children and residual scale $α$, the $L_0$-penalized reconstruction objective admits a closed-form phase boundary $λ_c(k,α)=α^2 k/(k-1)$: above it, pure parent absorption is strictly cheaper than pure child coding. Building on this boundary, we introduce HiPACE, an evaluation protocol that tests the boundary's structural consequence in real SAE dictionaries--measuring parent--child decoder structure over WordNet families, freezing the discovery-selected statistic before testing on unseen families, and contrasting genuine families against randomized sibling nulls. The boundary proves sharp in its native regime, predicting the synthetic transition within $\pm15%$ on all 30 tested cells. In Pythia-160m SAEs, the parent--child decoder gap recovers the predicted ordering with partial correlations up to $-0.93$ that sustain on the locked holdout and exclude sibling nulls ($p=0.002$). Controlled activation composition connects the theory's active-child count to the recovered family directions, and residual-stream interventions show that signed family directions increase parent-category logits, reversing under sign flip and vanishing under random controls--establishing causal sufficiency at the family-subspace level.

补充信息

↑