发表机构
Arizona State University(亚利桑那州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出可达费雪几何,分析参数高效微调中费雪迹对子群差距的揭示与局限,证明迹匹配可缩小差距但无法保证矩阵对齐,作为可扩展诊断工具而非严格证明。
AI 中文摘要
参数高效微调(PEFT)不仅决定了训练多少参数,还决定了模型能够移动的局部方向,因此相似的适配器可能对不同子群的损失产生不同影响。由于在适配器规模下构造曲率矩阵不可行,通常使用费雪迹等标量摘要。我们通过可达费雪(reachable Fisher)研究迹能揭示什么以及它丢失了什么:每个子群的全模型费雪通过适配器雅可比矩阵拉回。在似然损失下,它表示适配器可访问的高斯-牛顿曲率,其迹可通过得分梯度范数计算,无需构造完整矩阵。在匹配的子群梯度下,正定的可达费雪差,且其裕度超过海森-费雪缺陷,意味着任何足够小的非零模型改变更新都会增加符号差距。相比之下,受限算子范数决定最坏情况下的二次变化,而无矩阵的弗罗贝尼乌斯差异界定了其可达费雪分量。仅凭迹无法证明正定性或控制矩阵失配。相等的迹排除了正定差,但仍可能隐藏大的算子差异。在306个单种子模型中,较高的迹在1218次合格评估中的75.7%伴随更大的子群难度,而迹匹配在全部30个数据集-编码器-适配器组合中缩小了最佳-最差子群差距。然而,留出审计显示,算子差异在30个组合中的23个中减小,而无偏平方弗罗贝尼乌斯统计量仅在16个组合中减小。因此,费雪迹是可扩展的诊断和训练启发式方法,但不是局部差距行为或矩阵对齐的证明。
英文摘要
Parameter-efficient fine-tuning (PEFT) determines not only how many parameters are trained, but also which local directions a model can move in, so similar adapters can affect subgroup losses differently. Since curvature matrices are infeasible to form at adapter scale, scalar summaries such as the Fisher trace are often used instead. We study what the trace reveals and what it loses through the reachable Fisher: each subgroup's full-model Fisher pulled back through the adapter Jacobian. Under likelihood losses, it represents the Gauss-Newton curvature accessible to the adapter, and its trace can be computed from score-gradient norms without forming the full matrix. Under matched subgroup gradients, a positive-definite reachable-Fisher difference, with a margin exceeding the Hessian-Fisher defect, implies that every sufficiently small nonzero model-changing update increases the signed gap. In contrast, the restricted operator norm determines worst-case quadratic change, while a matrix-free Frobenius discrepancy bounds its reachable-Fisher component. Trace alone cannot certify definiteness or control matrix mismatch. Equal traces rule out a positive-definite difference but can still hide large operator discrepancies. Across 306 single-seed models, higher trace accompanies greater subgroup difficulty in 75.7 percent of 1,218 eligible evaluations, while trace matching reduces the best-worst subgroup gap in all 30 dataset-encoder-adapter combinations. However, held-out audits show that operator discrepancy decreases in 23 of 30 combinations, while the unbiased squared-Frobenius statistic decreases in only 16 of 30. Fisher trace is therefore a scalable diagnostic and training heuristic, but not a certificate of local gap behavior or matrix alignment.