arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

困惑度预测保护:在联邦参数高效微调中为最差客户端公平性选择预训练骨干网络

Perplexity Predicts Protection: Choosing Pretrained Backbones for Worst-Client Fairness in Federated Parameter-Efficient Fine-Tuning

Kiran Naseer, Samreen Azhar, Umar Shoaib, Haroon Mahmood, Muhammad Awais

arXiv 2609.23463首次发表:更新:

发表机构

University of Gujrat; Al Ain University; Qassim University(古吉拉特大学; 阿联酋大学; 卡西姆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过313次实验发现,在联邦LoRA微调中,较低困惑度的预训练骨干网络能显著提升最差客户端公平性,且个性化方法效果有限,建议在选择骨干前测量困惑度。

AI 中文摘要

联邦学习允许多方在不汇集数据的情况下训练共享模型,但当某个客户端的数据量远少于其他客户端时,即使群体的平均准确率看起来不错,该客户端也可能得到较差的服务。我们探究在LoRA微调下,预训练骨干网络的选择是否会影响这一情况,以及目标文本上的逐词困惑度能否在联邦训练开始前预测哪个骨干网络对最差客户端更有帮助。我们在三个文本分类数据集和三个规模相近的骨干网络(RoBERTa、BERTweet、PubMedBERT)上进行了313次实验,每个实验均与相同数据划分下的任务特定基线进行比较。较低困惑度的骨干网络持续为表现最差的客户端带来更大的提升,九个数据集-骨干网络组合的秩相关系数为-0.87;一个未参与分析的骨干网络也证实了这一模式。使用Ditto进行个性化仅恢复了单独训练与完全联邦之间差距的4-12%,而完全移除聚合则消除了这一益处。客户端的更新也未显示出与群体更新冲突的迹象;两者接近正交,排除了对这一失败的一种拟议解释。实际建议:在选择骨干网络前,先对任务文本样本测量困惑度,并且不要依赖个性化来保护数据匮乏的客户端。我们发布代码、预测和完整结果供他人测试。

英文摘要

Federated learning lets multiple parties train a shared model without pooling their data, but a client with far less data than the others can end up poorly served even when the group's average accuracy looks fine. We ask whether the choice of pretrained backbone affects this under LoRA fine-tuning, and whether per-word perplexity on the target text predicts which backbone helps the worst-off client before federated training starts. We ran 313 experiments across three text-classification datasets and three similarly sized backbones (RoBERTa, BERTweet, PubMedBERT), each compared against a task-specific baseline on identical data splits. Lower-perplexity backbones consistently produced larger gains for the worst-performing client, with a rank correlation of -0.87 across nine dataset-backbone pairs; a backbone held out of the analysis confirmed the pattern. Personalization with Ditto recovered only 4-12% of the gap between training alone and full federation, and removing aggregation entirely erased the benefit. A client's update also showed no sign of conflicting with the group's update; the two are close to orthogonal, ruling out one proposed explanation for this failure. Practically: measure perplexity on a sample of task text before choosing a backbone, and do not rely on personalization to protect a data-poor client. We release our code, predictions, and full results for others to test.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑