发表机构
Technical University Munich; Siemens AG(慕尼黑工业大学; 西门子公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究对比联邦预训练的下游微调(含全量、仅头部等变体)与GLUE文本下一个词预测的评估效果,发现后者与预训练测试困惑度相关性更强,提示仅下游微调易产生误导,需关注更接近预训练目标的评估信号。
AI 中文摘要
联邦预训练提供了一种在不集中底层数据集的情况下,基于私有或分布式数据训练基础模型的方法。然而,联邦预训练的评估仍具挑战性,因为客户端参与度和本地数据可用性的差异会导致直接可比的评估难以开展。此外,预训练测试困惑度与预训练分布相关,而下游基准引入的特定任务适配可能无法忠实地反映预训练期间建立的测试困惑度。本研究旨在探究哪种评估协议能更可靠地反映联邦预训练的质量。我们使用一组受控的中心化和联邦训练的模型,这些模型是在相同客户端数据上训练的16M参数Transformer模型,通过评估协议是否保留在相同预训练测试集上建立的参考排名来评估这些协议。我们比较了在GLUE上的下游微调,包括全量、仅头部和减少数据的变体,以及在GLUE文本上的下一个词预测作为内在评估信号。结果显示,下游微调无法可靠地保留预训练排名,而直接下一个词预测与预训练测试困惑度具有强相关性。这些发现表明,在比较联邦预训练模型时,仅下游微调可能会产生误导,更应关注与原始预训练目标更接近的评估信号。
英文摘要
Federated pre-training offers a way to train foundation models on private or distributed data without centralizing the underlying datasets. However, evaluating federated pre-training remains challenging because differences in client participation and local data availability can make directly comparable evaluation difficult. Moreover, pre-training test perplexity is tied to the pre-training distribution, while downstream benchmarks introduce task-specific adaptation that may not faithfully reflect the test perplexity established during pre-training. In this work, we study which evaluation protocol more reliably reflects federated pre-training quality. Using a controlled set of centralized and federated-trained models of a 16M parameter transformer model trained on identical client data, we assess evaluation protocols by whether they preserve a reference ranking established on the same pre-training testset. We compare downstream fine-tuning on GLUE, including full, head-only, and reduced-data variants, with next-token prediction on GLUE text as an intrinsic evaluation signal. Our results show that downstream fine-tuning does not reliably preserve the pre-training ranking, whereas direct next-token prediction exhibits a strong correspondence with the pre-training test perplexity. These findings suggest that downstream fine-tuning alone can be misleading when comparing federated pre-trained models, and that evaluation signals closer to the original pre-training objective deserve greater attention.