arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27593cs.LGcs.AI

隐藏而非删除:网络如何抑制纠缠特征

Hidden not Deleted: How Networks Suppress Entangled Features

Akash Samanta, Manish Pratap Singh, Debasis Chaudhuri

首次发表
浏览论文内容

中文总结 AI 辅助

本文揭示线性擦除在密集叠加下失效,网络通过镜像或阴影解非线性抑制特征,留下可恢复痕迹,解释了LLM遗忘中抑制而非删除的失败机制。

中文摘要 AI 辅助

通过线性投影进行的概念擦除方法假设特征占据可分离的子空间。我们证明这一假设在密集叠加下失效:当两个特征被迫进入共享单一子空间的对映对时,最先进的线性擦除会同时破坏两者,而不仅仅是目标特征。经过梯度下降训练的网络反而以非线性方式解决此问题,但并非均匀解决:它们根据初始化收敛到两种不同的电路级解决方案之一,我们称之为镜像解和阴影解。我们将这种分岔映射为特征纠缠的函数,表明它反映了稳定的吸引子结构而非我们设置的伪影,并使用有针对性的因果干预证明两种解决方案都保留了被擦除特征表示的实质性、可测量的痕迹,可通过单个标量补丁恢复,而无需任何进一步训练。这镜像了最近在LLM遗忘中经验观察到的失败模式,其中抑制而非删除使被遗忘的知识重新浮现;我们的结果为该失败模式为何发生提供了机制性的、因果验证的解释。

英文摘要

Concept erasure methods that operate via linear projection assume that features occupy separable subspaces. We show this assumption fails under dense superposition: when two features are forced into an antipodal pair sharing a single subspace, state-of-the-art linear erasure destroys both, not just the target. Networks trained with gradient descent instead solve this problem non-linearly, but not uniformly: they converge to one of two distinct circuit-level solutions depending on initialization, which we call mirror and shadow solutions. We map this bifurcation as a function of feature entanglement, show it reflects a stable attractor structure rather than an artifact of our setup, and use targeted causal interventions to demonstrate that both solutions leave a substantial, measurable trace of the erased feature's representation intact, recoverable through a single scalar patch rather than requiring any further training. This mirrors a failure mode recently observed empirically in LLM unlearning, where suppression rather than deletion allows forgotten knowledge to resurface; our results offer a mechanistic, causally-validated account of why that failure mode occurs.

发表机构

  • Techno India University(泰克诺印度大学)
  • DRDO Young Scientist Laboratory - CT(DRDO青年科学家实验室-CT)

机构由 AI 辅助整理,请以论文原文为准。

↑