发表机构
Dalhousie University; Johns Hopkins University; Indian Institute of Management Bangalore; Vector Institute(达尔豪斯大学; 约翰斯·霍普金斯大学; 印度管理研究所班加罗尔分校; 向量研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文证明仅凭发布时的互信息无法认证开放权重模型的防篡改能力,需约束攻击动态,并通过构造和实验验证了顺序与参数化依赖的恢复速度。
AI 中文摘要
移除有害信息是否能使开放权重模型抵抗微调攻击?我们表明,发布时的互信息单独无法普遍证明恢复速度缓慢。保持功能的重新参数化在改变梯度下降几何结构的同时保持信息不变,因此不变性证书受最快可达参数化的限制。我们将这一原理应用于训练数据过滤下的权重-数据互信息和能力移除下的标签-表示互信息。在固定权重-数据信息下,训练顺序可以改变恢复时间,而精确的表示级独立性可以保留整个参数雅可比矩阵。一个显式构造使两个信息量都等于零,并在一个梯度步骤中恢复。受控实验说明了顺序依赖的恢复和参数化依赖的攻击速度。这些结果确定了认证所缺失的要求:对发布时互信息之外的攻击动态的约束。
英文摘要
Does removing harmful information make open-weight models resistant to fine-tuning attacks? We show that mutual information at release alone cannot universally certify slow recovery. Function-preserving reparameterizations leave information unchanged while altering gradient-descent geometry, so an invariant certificate is bounded by the fastest reachable parameterization. We apply this principle to weight--data mutual information under training-data filtering and label--representation mutual information under capability removal. Training order can change recovery time at fixed weight--data information, while exact representation-level independence can preserve the entire parameter Jacobian. An explicit construction has both information quantities equal to zero and recovers in one gradient step. Controlled experiments illustrate order-dependent recovery and parameterization-dependent attack speed. These results identify the missing requirement for certification: constraints on attack dynamics beyond mutual information at release.
CommentsUnder submission AISTATS 2026