arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

神经表示中的认证拓扑交互:类别解缠主要是成对性的

Certified Topological Interaction in Neural Representations: Exact Tests and the Statistic They Require

Sushovan Majhi

arXiv 2609.08561首次发表:更新:

发表机构

George Washington University(乔治华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出用交点欧拉特征轮廓认证神经表示中的类别解缠,发现解缠主要是成对性的,且数据增强是唯一有效分离类别的训练选择。

AI 中文摘要

类别解缠(表示中类别条件点云沿深度和训练过程中的分离)通常通过描述性曲线来解读。我们使用最近引入的交点欧拉特征轮廓,将其度量为由标记点云之间的认证拓扑交互:即点云球并集重叠部分的欧拉特征作为尺度的函数,通过一次Alpha复形扫描计算,无需边界矩阵约简。每个数值都附带一项检验:两个方向的精确置换检验、一个受保护的分离证书,以及用于比较性应用主张的配对检验。在111个训练网络和52,650个认证测量中,解缠具有深度分级性,并集中在最初几个训练周期内,交互商数按混淆度对类别对进行排序(Spearman rho=0.83),与廉价的可分性统计相当。在一个96模型的因子设计中,数据增强是唯一能相对于随机水平分离类别的训练选择;权重衰减压缩重叠但不分离,而深度和宽度无影响。这一结构性发现只有k阶统计量才能提出:在三重类别联合纠缠度低于其最强对的纠缠度,在97%的三重层单元和99.5%的深层单元中如此,远低于测量的零假设下限,在视觉编码器和冻结语言模型中均如此。这种成对优势是一种规律而非定律:从重叠的嵌套性可预期,但并非由几何强制决定,在初始化和原始像素中存在,并在网络记忆随机标签时仅在最后阶段产生。未归一化的轮廓质量预测测试准确率(R^2=0.94),而商数则不能,且两者都不优于线性探针。一个教训被完整报告:配对检验必须使用无尺度统计量,否则会将特征范数动态认证为解缠。

英文摘要

Class disentanglement--the separation of a representation's class-conditional point clouds along depth and over training--is measured by descriptive curves: the sentence such a study wants to write, layer l+1 is more disentangled than layer l, is an eyeball judgement with no null. We supply the inferential layer for a topological measurement of class overlap, the Intersection Euler Characteristic Profile: the Euler characteristic of the overlap of the clouds' ball unions as a function of scale, from one Alpha-complex sweep with no boundary-matrix reduction. Every number carries a test--exact permutation tests in both directions, a guarded separation certificate the invariant requires, and a paired sign-flip test for comparative claims. Building that test taught a lesson outliving this invariant: its statistic must be scale-free. On the raw profile mass, which has units of feature length, 12,375 paired tests return 5,633 significant steps of which every one at the first epoch points the wrong way, certifying feature-norm dynamics as disentanglement; the dimensionless statistic returns 2,707, with 2,026 decreases. Across 111 networks and 52,650 measurements, disentanglement is depth-graded and early, and interaction quotients rank class pairs by confusability (rho=0.83), on par with cheap separability statistics. In a 96-model factorial, augmentation is the one training choice that separates classes relative to chance; weight decay compresses the overlap without separating. Only a k-fold statistic can pose the structural question: the joint entanglement of a class triple sits below its strongest pair in 97% of triple-layer cells and 99.5% of deep cells, at median ratios far below a measured null floor, in vision encoders and frozen language models--a regularity, not a law. The unnormalized mass predicts test accuracy (R^2=0.94), the quotient does not, and neither beats a linear probe.

Comments36 pages, 9 figures, 4 tables. Code and measurement records: https://github.com/sushovan4/disentanglement

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑