发表机构
RPTU Kaiserslautern-Landau; University of the Witwatersrand(莱茵兰-普法尔茨凯泽斯劳滕工业大学; 威特沃特斯兰德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出拓扑概念表示(TCR),一种事后操作抽象,通过拓扑编码概念组织,统一干预框架,提升擦除方法最差组准确率21.89和概念迁移可恢复性5.54。
AI 中文摘要
基于概念的方法为解释和操作学习到的表示提供了语义层面,但现有的编辑方法通常专门针对特定干预,且不提供概念组织的通用可编辑表示。为实现这一目标,我们引入了拓扑概念表示(TCR),这是一种事后操作抽象,共同表征学习到的表示中编码的概念及其关系。TCR从概念可恢复性和交互分数构建中间概念空间,并通过拓扑紧凑地编码其组织。干预通过修改此抽象来表达,并将修改传播回底层学习到的表示。这允许不同的概念级操作共享同一优化框架,并将期望的概念组织与实现机制分离。我们建立了TCR的稳定性和重参数化不变性属性,及其与现有概念编辑公式的联系。我们使用TCR作为现有擦除方法的预处理步骤来解缠概念,在可比概念泄漏下平均将最差组准确率提高21.89。我们进一步使用TCR将概念从教师模型转移到学生模型,将概念可恢复性提高最多5.54,同时提高或保持具有竞争力的测试top-1准确率。
英文摘要
Concept-based methods provide a semantic level for interpreting and manipulating learned representations, but existing editing approaches are typically specialized to particular interventions and do not provide a common and editable representation of concept organization. To achieve this, we introduce Topological Concept Representations (TCR), a post-hoc operational abstraction that jointly characterizes the concepts encoded in a learned representation and their relationships. TCR constructs an intermediate concept space from concept recoverability and interaction scores, and compactly encodes its organization through topology. Interventions are expressed through modifications of this abstraction that are propagated back to the underlying learned representation. This allows different concept-level operations to share the same optimization framework and separates the desired concept organization from the mechanism to achieve it. We establish stability and reparameterization-invariance properties of TCR and its connections to existing concept-editing formulations. We use TCR to disentangle concepts as a preprocessing step for existing erasure methods, improving worst-group accuracy by 21.89 on average at comparable concept leakage. We further use TCR to transfer concepts from teacher to student models, improving concept recoverability by up to 5.54 while also improving or maintaining competitive test top-1 accuracy.