arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

预训练塑造谱结构:基础模型中OOD鲁棒性的架构与策略条件预测

Pretraining Shapes Spectral Structure: Architecture- and Strategy-Conditional Prediction of OOD Robustness in Foundation Models

Sangyoon Bae, Sk Miraj Ahmed, Shinjae Yoo, Jiook Cha

arXiv 2610.09709首次发表:更新:

发表机构

Seoul National University; Brookhaven National Laboratory(首尔大学; 布鲁克海文国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究证明基础模型的OOD鲁棒性可由预训练权重的谱结构预测,提出架构与策略条件统计量,在116个模型上验证,单元内排序准确率达92%,并缩小OOD差距24%。

AI 中文摘要

我们能否在获得任何目标数据之前,就确定一个基础模型是否会泛化到分布外(OOD)场景?现有的诊断方法需要源数据或目标数据,这在目标域存在之前就排除了它们的适用性。那些仅使用权重的方法则对每种架构应用同一种统计量,无法区分鲁棒模型与脆弱模型。我们证明,答案编码在预训练权重的谱结构中。两种力量塑造了这种结构。架构决定了信息如何在权重矩阵中存储。预训练策略决定了什么被奖励。它们共同设定了一种谱几何,该几何支配着OOD鲁棒性。我们证明了OOD准确率差距受源表示集中程度的紧密程度限制。仅从预训练权重计算出的统计量可作为该集中程度的代理。该代理的方向在架构族之间发生反转。我们将其操作化:该方向在(架构×策略)组合内是稳定的,这是我们测试的最细粒度分组,称为单元。汇总跨越7种模态的116个模型,单一统计量对OOD鲁棒性的排序较弱,因为相反方向的单元相互抵消。在单元内,为其选择的统计量对92%的模型对按OOD鲁棒性进行了样本内排序。该选择不泄露目标:对于矩阵之外的每个模型族,我们在运行其OOD评估之前记录了单元、指标和符号,预测方向在每种情况下都成立:脑电图、基因组和蛋白质。基于谱集中度采取行动,在87.5%的ID保留率下将OOD差距缩小了24%。该诊断仅基于已发布的权重运行,因此OOD鲁棒性在模型选择时即可检查,在数据和计算资源投入目标域之前。

英文摘要

Can we determine whether a foundation model will generalize out-of-distribution (OOD) before any target data is available? Existing diagnostics require source or target data, which rules them out before a target domain exists. Those that use the weights alone apply one statistic to every architecture, and do not separate robust models from fragile ones. We show the answer is encoded in the spectral structure of pretrained weights. Two forces shape that structure. Architecture determines how information is stored in weight matrices. Pretraining strategy determines what is rewarded. Together they set a spectral geometry that governs OOD robustness. We prove that the OOD accuracy gap is bounded by how tightly the source representations concentrate. A statistic computed from the pretrained weights alone serves as a proxy for that concentration. The direction of that proxy reverses between architecture families. We operationalize it: the direction is stable within one (architecture X strategy) combination, the finest grouping we test, which we call a cell. Pooled over 116 models spanning 7 modalities, a single statistic ranks OOD robustness weakly, because cells of opposite direction cancel. Within a cell, the statistic selected for it orders 92% of model pairs by OOD robustness in-sample. The selection does not leak the target: for each model family outside the matrix we logged the cell, metric and sign before running its OOD evaluation, and the predicted direction held in every case: EEG, genomic and protein. Acting on spectral concentration narrows the OOD gap by 24% at 87.5% ID retention. The diagnostic operates on released weights alone, so OOD robustness becomes checkable at model-selection time, before data or compute is committed to a target domain.

Comments10 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑