发表机构
The University of Sydney; UNSW Sydney(悉尼大学; 新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出更新压力密度动力学解释训练-验证分离现象,通过条件局部模型和固定训练探针揭示梯度分配异质性,并在ResMLP、NLP模型及ResNet-18上验证,发现私有特征占比增加会扩大准确率差距。
AI 中文摘要
训练-验证分离是指在已观察的训练示例与有限的留出验证集上性能之间不断演化的差异。我们提出了一种关于预训练模型适应过程中这种差距如何发展的动态结构解释:持续拟合可以将更新需求从广泛可复用的支持转向支持范围较窄且留出迁移能力较弱的区域。一个条件局部模型将这种转变与梯度分配异质性的增加以及训练-验证分离联系起来。固定训练探针使得这种结构演化在没有验证示例进入读出器的情况下可被观察;留出性能被单独用于评估其与差距的关系。在一个由残差多层感知器(ResMLP)实现的构造层次结构中,将示例私有特征的目标占比从$p=.3$增加到$.5$再到$.7$,同时保持四个共享特征层级之间的相对混合比例$1{:}2{:}3{:}4$不变,每个条件下五次运行的平均最终准确率差距从$.185$增加到$.331$再到$.527$。分别在训练和验证示例上测量的掩码输入损失暴露了相应的迁移不对称性。自然语言处理(NLP)分析使用RoBERTa、DeBERTa和Qwen在六个数据集上进行10轮运行(共90次运行):训练探针加权的类内和整体离散度读出器与准确率差距的原始和平滑水平相关性在全部90次运行中均为正。大约一个epoch间隔配对的原始变化分别在86/90和87/90次运行中保持正相关。一项40轮的ResNet-18研究在三个视觉数据集上测试了两种读出器。总之,受控模拟、NLP和视觉实验在多种设置下支持动态结构解释,真实模型证据在指定监控器下检验了其可观察的预测。
英文摘要
Train-validation separation is the evolving difference between performance on observed training examples and a finite held-out validation set. We propose a dynamic structural account of how this gap develops during adaptation of pretrained models: continued fitting can shift update demand from broadly reusable support toward narrower support with weaker held-out transfer. A conditional local model links this shift to increasing heterogeneity in gradient allocation and train-validation separation. Fixed training probes make this structural evolution observable without validation examples entering the readouts; held-out performance is used separately to evaluate its relation to the gap. In a constructed hierarchy implemented with a residual multilayer perceptron (ResMLP), increasing the target share of example-private features from $p=.3$ to $.5$ to $.7$, while preserving the relative mixture $1{:}2{:}3{:}4$ among the four shared feature levels, increases the final mean accuracy gap from $.185$ to $.331$ to $.527$ across five runs per condition. Masked-input losses measured separately on training and validation examples expose the corresponding transfer asymmetry. The natural language processing (NLP) analysis uses 10-epoch runs of RoBERTa, DeBERTa, and Qwen on six datasets (90 runs): the training-probe-weighted within-class and overall dispersion readouts each have positive raw and smoothed level correlations with the accuracy gap in all 90 runs. Raw changes paired at approximately one-epoch intervals remain positively associated in 86/90 and 87/90 runs, respectively. A 40-epoch ResNet-18 study tests both readouts on three vision datasets. Together, controlled simulation, NLP, and vision support the dynamic structural account across settings, with real-model evidence testing its observable predictions under the specified monitors.