arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18736cs.LGcs.CRcs.DC

FedLNS:利用层归一化签名建模缓解联邦大语言模型中的对抗性操纵

FedLNS: Leverage LayerNorm Signature Modeling to Mitigate Adversarial Manipulation in Federated LLMs

  • Interdisciplinary Centre for Security, Reliability and Trust (SnT), University of Luxembourg(卢森堡大学安全、可靠性与跨学科研究中心(SnT))
  • Carnegie Mellon University(卡内基梅隆大学)
  • Edith Cowan University(埃迪斯科文大学)
  • TU Berlin(柏林工业大学)

机构由 AI 辅助整理,请以论文原文为准。

Kai Li, Jong-Ik Park, Carlee Joe-Wong, Wei Ni, Falko Dressler

AI总结:

FedLNS是一种服务器端联邦学习框架,通过层归一化签名筛选恶意更新,在200个客户端、40%目标操纵下,对三类模型均实现优于基线的测试困惑度。

AI中文摘要:

联邦训练使语言模型能够从分布式私有文本中学习,但服务器无法直接验证产生每个客户端更新的本地监督或优化过程。因此,恶意客户端可以在损坏的目标上训练,引入不正确的上下文-标记关联,并通过重复聚合降低全局模型的性能,这种性能下降还会增加生成不可靠或幻觉内容的风险。我们提出了联邦学习归一化签名框架FedLNS,这是一种用于轻量级恶意更新筛选的服务器端框架。FedLNS通过可训练归一化层参数的变化来表示每个客户端更新,并针对稳健的、感知历史的跨客户端参考来筛选可疑更新。由于签名是从服务器返回的本地模型中提取的,与标准联邦学习(FL)方法相比,FedLNS不需要额外的客户端到服务器参数或元数据交换。筛选后,保留的全模型更新可以使用标准FL或其他兼容的聚合规则进行聚合。FedLNS不需要原始客户端数据、可信服务器数据集、标记的攻击示例或单独训练的检测器。在200个客户端从头开始训练的GPT风格、BERT风格和LLaMA风格模型上进行的实验表明,在40%的总体目标操纵下,FedLNS在IID(独立同分布)和非IID数据划分下,对于所有三种架构,都比六个基线中最强的一个实现了更低的测试困惑度。

英文摘要:

Federated training enables language models to learn from distributed private text, but the server cannot directly verify the local supervision or optimization process that produces each client update. A malicious client can therefore train on corrupted targets, introduce incorrect context-token associations, and degrade the global model through repeated aggregation. Such degradation can also increase the risk of unreliable or hallucinatory generation. We propose Federated Learning with Normalization Signatures (FedLNS), a server-side framework for lightweight malicious-update screening. FedLNS represents each client update through changes in trainable normalization-layer parameters and screens suspicious updates against a robust, history-aware cross-client reference. Because the signatures are extracted at the server from the returned local models, FedLNS requires no additional client-to-server parameter or metadata exchange compared to standard federated learning (FL) methods. After screening, the retained full-model updates can be aggregated using standard FL or another compatible aggregation rule. FedLNS requires no raw client data, trusted server dataset, labeled attack examples, or separately trained detector. Experiments on GPT-style, BERT-style, and LLaMA-style models trained from scratch with 200 clients show that, under 40% population-level target manipulation, FedLNS achieves lower test perplexity than the strongest of six baselines for all three architectures under both IID (independently and identically distributed) and non-IID data partitions.

补充信息

↑