arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35351cs.LG

概念提取中超越几何的干扰

Interference Beyond Geometry in Concept Extraction

发表机构阿尔伯塔大学 · 阿尔伯塔机器智能研究所 · 哈佛大学
查看机构详情
  • University of Alberta(阿尔伯塔大学)
  • Amii(阿尔伯塔机器智能研究所)
  • Harvard University(哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

Valérie Costa, Bahareh Tolooshams

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出有效干扰概念,结合特征几何与代码统计,刻画架构约束通过四种机制影响干扰,实验表明干扰取决于特征使用及架构,超越单纯几何。

中文摘要 AI 辅助

干扰通常被视为学习特征之间的几何重叠。我们引入了有效干扰,它结合了特征几何和代码统计来捕捉实际发生的交互,区分了建设性干扰与破坏性干扰,以及频繁的弱交互与罕见的强交互。在局部固定支撑假设下,我们刻画了架构约束如何通过四种机制塑造干扰:特征正交化、偏置补偿、增益适应和编码器-解码器分离。使用稀疏自编码器的实验表明,受约束的架构选择性地减少共激活特征之间的重叠,而偏置、增益和编码器自由度允许建设性的交叉贡献得以保留。总之,这些结果表明,学习表示中的干扰不仅取决于特征几何,还取决于特征的使用方式以及产生其代码的架构。

英文摘要

Interference is commonly treated as geometric overlap between learned features. We introduce effective interference, which combines feature geometry and code statistics to capture realized interactions, distinguishing constructive from destructive interference and frequent weak interactions from rare strong ones. Under local fixed-support assumptions, we characterize how architectural constraints shape interference through four mechanisms: feature orthogonalization, bias compensation, gain adaptation, and encoder-decoder separation. Experiments with sparse autoencoders show that constrained architectures selectively reduce overlap among co-active features, while bias, gain, and encoder freedom allow constructive cross-contributions to remain. Together, these results show that interference in learned representations depends not only on feature geometry, but also on how features are used and on the architecture that produces their codes.

↑