迭代擦除计数不是仿射不变的概念维度
Iterative Erasure Count Is Not an Affine-Invariant Concept Dimension
浏览论文内容
中文总结 AI 辅助
该研究指出迭代擦除计数不是仿射不变的概念维度,其为与过程相关的估计量,而非本质语义维度,且通过多种实验验证了可逆重参数化会改变迭代擦除相关计数。
中文摘要 AI 辅助
神经表征用多少个方向来编码一个概念?一个常见的答案是反复擦除探测方向并报告停止计数或累计移除秩。我们证明,在保信息的可逆重参数化下,这两个量都会发生变化,因此它们都不是本质的概念维度。我们将模型定义的总体量(生成维度、充分线性维度和最小保护秩)与过程定义的量(如停止计数和累计编辑秩)区分开来。在总体高斯构造中,可逆剪切保留了预测问题和所有三个量,但将累计欧几里得擦除计数从1变为2。对于Moore–Penrose普通最小二乘以及每个有限非负岭权重,这种分离都成立。对于匹配我们的动机视频分析的双输出全-QR过程,累计编辑秩同样从2变为环境维度4。相反,当正定度量、探测、正则化器和打破平局的方式被一致迁移时,完整的累计度量-QR轨迹是仿射等变的;精确协方差是其推论,而非规范语义度量。在已知秩的有限样本Adam/QR校准中,在全部20次大样本运行中,恒等混合在一次接受更新后停止,而每个测试的剪切$a\in\{.5,.75,1,1.25,2\}$在全部20次运行中都至少接受两次更新。对冻结V-JEPA2特征的可控重参数化保留了零秩预测,但改变了实际优化下的后续欧几里得轨迹。这些视觉接触实验是压力测试,而非接触维度的估计。因此,迭代擦除返回一个由表征几何和完整测量过程共同决定的、与过程相关的估计量,而非本身的语义维度。
英文摘要
How many directions does a neural representation use to encode a concept? A common answer repeatedly erases probe directions and reports the stopping count or cumulative removed rank. We show that both quantities can change under an information-preserving invertible reparameterization, so neither is intrinsically a concept dimension. We distinguish model-defined population quantities (generating dimension, sufficient linear dimension, and minimum guarding rank) from procedure-defined quantities such as stopping count and cumulative edit rank. In a population Gaussian construction, an invertible shear preserves the prediction problem and all three quantities, yet changes the cumulative Euclidean erasure count from one to two. The separation holds for Moore--Penrose ordinary least squares and every finite nonnegative ridge weight. For a two-output full-QR procedure matching our motivating video analysis, cumulative edit rank similarly changes from two to the ambient dimension four. Conversely, the complete cumulative metric-QR trajectory is affine-equivariant when its positive-definite metric, probe, regularizer, and tie-breaking are transported consistently; exact covariance is one corollary, not a canonical semantic metric. In a known-rank finite-sample Adam/QR calibration, identity mixing stops after one accepted update in all 20 large-sample runs, whereas each tested shear $a\in\{.5,.75,1,1.25,2\}$ accepts at least two updates in all 20 runs. Controlled reparameterizations of frozen V-JEPA2 features preserve rank-zero predictions yet alter later Euclidean trajectories under practical optimization. These visual contact experiments are stress tests, not estimates of contact dimension. Iterative erasure therefore returns a procedure-relative estimand jointly determined by representation geometry and the full measurement procedure, not a semantic dimension by itself.
发表机构
- UCLA(加州大学洛杉矶分校)
- Cal Poly Pomona(加州州立理工大学波莫纳分校)
机构由 AI 辅助整理,请以论文原文为准。