AI 中文总结
本研究通过对抗训练作为受控工具,首次实证检验表征简单性是否转化为电路规模简单性,发现二者在阈值依赖方式上分离,鲁棒模型在高保真度下电路更小。
AI 中文摘要
稀疏自编码器的可分解性和集中特征归因日益被视为模型计算更易于逆向工程的证据。这种表征和归因上的清晰性是否实际上预测了更小或更易处理的因果电路,仍是一个开放问题。我们使用对抗训练作为受控工具直接测试这一点:它可靠地重塑内部表征,但仅此并不构成对电路规模的测试。我们通过逆向工程复杂性来研究这个问题:在固定保真度水平下恢复模型行为所需的因果结构。据我们所知,这是首次对表征或归因简单性是否转化为电路层面的因果简单性进行的受控实证测试。从相同的预训练GPT-2 Small检查点出发,我们应用匹配的标准和对抗持续训练,要求两种条件在间接宾语识别上保持能力并通过独立的鲁棒性验证,然后比较机制。随后,我们沿三个互补轴比较模型:稀疏自编码器可分解性、SAE特征在任务归因中的参与度,以及从原始计算图恢复的忠实电路规模。鲁棒模型更具SAE可分解性,并在任务归因中涉及更少的SAE特征。电路规模具有依赖性:在能力匹配的IOI上,标准模型在低于85%保真度时领先或持平,但鲁棒模型在高保真度(90%、95%)下需要显著更少的边,这一模式在主要配对中确立,而表征趋势在七点扫描和第二个语料库中普遍存在。
英文摘要
Sparse-autoencoder decomposability and concentrated feature attribution are increasingly treated as evidence that a model's computation is easier to reverse-engineer. Whether this representational and attributional cleanliness actually predicts a smaller or more tractable causal circuit remains an open question. We test this directly using adversarial training as a controlled instrument: it reliably reshapes internal representations, but this alone does not constitute a test of circuit size. We investigate this question through reverse-engineering complexity: the causal structure required to recover a model's behavior at a fixed level of faithfulness. To our knowledge, this is the first controlled empirical test of whether representational or attributional simplicity translates into causal simplicity at the circuit level. Starting from the same pretrained GPT-2 Small checkpoint, we apply matched standard and adversarial continual training, requiring both conditions to retain competence on indirect object identification and pass independent robustness verification before comparing mechanisms. We then compare the models along three complementary axes: sparse-autoencoder decomposability, SAE feature engagement in task attribution, and the size of faithful circuits recovered from the raw computational graph. The robust model is more SAE-decomposable and engages fewer SAE features in task attribution. Circuit size is regime-dependent: on competence-matched IOI, standard leads or ties below 85% faithfulness, but robust needs substantially fewer edges at high faithfulness (90%, 95%), a pattern established on the primary pair while representational trends generalize across a seven-point sweep and a second corpus.
CommentsUnder review at ICLR 2027. 21 pages, 5 figures, 12 tables