发表机构
Modulabs(Modulabs)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种JEPA训练配方,使潜在项在表格基础模型中与值目标共存,实验表明其性能略逊于仅值目标,但揭示了固定步数预算的混淆效应。
AI 中文摘要
表格基础模型学习在上下文中预测单元格值,而世界模型的自监督学习要求在表示空间中进行预测(LeCun, 2022; Assran et al., 2023)。在表格基础模型先验上,联合嵌入预测架构(JEPA)的潜在项在我们早期的运行中崩溃,并将编码器一同拖入常数映射。我们报告了一种配方,在该配方下,潜在项在收敛时与值目标共存:值头读取编码器字段而非预测器,目标是指数移动平均(EMA)差异。为了限制其相对于仅值目标的成本,两个分支都训练直到平台规则停止它们,没有固定的步数预算。固定的训练时长曾将性能放缓误认为性能上限,因为仅值分支在通常预算之后仍在显著改进。在收敛时,每个分支各运行一次,JEPA分支在147个真实数据集上落后于仅值分支,分类任务上胜负为32:70(每个数据集名称一条记录时为29:63),回归任务上为8:24,分类上的差距较小,回归上的差距较大,且在每个分层和每个基准中计数趋势相同。JEPA分支(jepa)达到平台所需的步数是仅值分支(ds)的1.42倍,墙钟时间是后者的1.66倍。
英文摘要
Tabular foundation models learn to predict cell values in context, whereas world-model self-supervision asks for prediction in representation space (LeCun, 2022; Assran et al., 2023). On a tabular foundation-model prior, the latent term of a joint-embedding predictive architecture (JEPA) collapsed in our earlier runs and took the encoder with it to a constant map. We report a recipe under which the latent term survives to convergence beside the value objective: the value head reads the encoder field rather than the predictor, and the target is an exponential moving average (EMA) difference. To bound its cost against the value-only arm, both arms train until a plateau rule stops them, with no fixed step budget. A fixed horizon had confounded a slowdown with a ceiling, since the value-only arm was still improving well past the usual budget. At convergence, in one run per arm, the JEPA arm trails the value-only arm across 147 real datasets, 32:70 wins to losses on classification (29:63 with one entry per dataset name) and 8:24 on regression, the margin small on classification and wider on regression, and the count leans the same way in each stratum and each benchmark. The JEPA arm (jepa) needs 1.42 times as many steps as the value-only arm (ds), and 1.66 times its wall-clock, to reach its plateau.
Comments16 pages, 5 figures