Superposition as Lossy Compression: Measure with Sparse Autoencoders and Connect to Adversarial Vulnerability
叠加作为损失性压缩:通过稀疏自编码器测量并连接到对抗脆弱性
机构 * University of Amsterdam(阿姆斯特丹大学) ; Toronto Metropolitan University(多伦多 Metropolitan 大学) ; Vector Institute for Artificial Intelligence(人工智能向量研究所)
AI总结 本研究通过信息论框架测量神经网络的叠加程度,揭示其与对抗鲁棒性的关系,表明叠加在不同任务复杂性和网络容量下呈现不同表现。
Comments Accepted to TMLR, view HTML here: https://leonardbereska.github.io/blog/2025/superposition/
Journal ref Transactions on Machine Learning Research, 2025