arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

特征叠加中线性可达性的高概率保证

High-probability guarantees for linear accessibility in feature superposition

Enrico Vompa

arXiv 2609.09556首次发表:更新:

发表机构

Tallinn University of Technology(塔林理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过压缩感知框架推导特征叠加中线性可达性的高概率维度界,证明其线性增长而非二次,并验证了该界,为稀疏自编码器评估和神经可解释性提供理论框架。

AI 中文摘要

神经网络可以利用特征叠加来编码比维度更多的概念,但跨特征干扰限制了同时活跃特征的线性可达性。通过将线性可达性视为压缩感知问题,我们在次高斯噪声下推导了固定支撑集的高概率界,证明了充分维度呈线性增长($d=O_{\varepsilon}(k \log m)$),而非先前最坏情况下的二次方极限。随后,我们通过高斯尾部近似在系统参数范围内验证了这些界。这些结果量化了线性表示假设的几何约束,为评估稀疏自编码器、组合泛化和神经可解释性提供了框架。

英文摘要

Neural networks can leverage feature superposition to encode more concepts than dimensions, but cross-feature interference constrains the linear accessibility of simultaneously active features. By framing linear accessibility as a compressed sensing problem, we derive high-probability bounds for fixed supports under subgaussian noise, proving the sufficient dimension scales linearly ($d=O_{\varepsilon}(k \log m)$) rather than prior worst-case quadratic limits. We characterize the asymmetry between active and inactive interference and the trade-off between interference and observation-noise budgets. We then validate these bounds across system parameters through Gaussian-tail approximations. We also introduce IHT-SAE, which uses learned iterative refinement to improve feature recovery beyond the limits of linear availability. These results quantify the geometric constraints of the linear representation hypothesis, providing a framework for evaluating sparse autoencoders, compositional generalization, and neural interpretability.

Commentspreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑