发表机构
Università della Svizzera italiana (USI); MIT Media Lab, Massachusetts Institute of Technology(意大利语瑞士大学; 麻省理工学院媒体实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出分布式潜意识学习(DSL),通过传输任务无关输入的载体输出而非模型更新,实现协作学习,在LLM组合和联邦分类中显著提升性能并降低通信开销。
AI 中文摘要
协作学习通常交换模型参数:联邦客户端通信更新,而独立适应的基础模型通过交换适配器或检查点进行组合。这使得通信规模随模型大小增长,并需要在权重空间中调和局部特化,而权重空间中干扰常见。我们探究知识是否可以通过模型在任务无关输入上的行为来共享。我们引入了分布式潜意识学习(DSL),一种协作学习原语,其中参与者本地适应共同模型,用任务无关输入探测它,并仅传输产生的载体输出。协调器汇集这些输出并将其蒸馏到共享模型中。该原语通过载体补全支持一次性基础模型组合,通过载体逻辑支持迭代联邦学习,无需传输模型更新或需要任务相关的代理数据。在LLM组合中,与LoRA平均相比,DSL实现了更高的偏好保留率(94.56%对87.76%)和更大的GSM8K增益(22.0对0.6点),同时减少了30.6-49.0倍的上传量。在联邦分类中,DSL在MNIST上达到96.83%的准确率,上行链路比FedAvg少8.9倍,并在CIFAR-10和Tiny ImageNet上提供了更低通信的操作点。这些结果确立了随机载体输出作为跨不同协作学习范式知识共享的实用通信原语。
英文摘要
Collaborative learning typically exchanges model parameters: federated clients communicate updates, while independently adapted foundation models are combined by exchanging adapters or checkpoints. This makes communication scale with model size and requires local specializations to be reconciled in weight space, where interference is common. We ask whether knowledge can instead be shared through model behavior on task-unrelated inputs. We introduce Distributed Subliminal Learning (DSL), a collaborative learning primitive in which participants adapt a common model locally, probe it with task-unrelated inputs, and transmit only the resulting carrier outputs. A coordinator pools these outputs and distills them into a shared model. The primitive supports one-shot foundation-model composition through carrier completions and iterative federated learning through carrier logits, without transmitting model updates or requiring task-related proxy data. In LLM composition, compared with LoRA averaging, DSL achieves higher preference retention (94.56% vs. 87.76%) and a larger GSM8K gain over the base model (22.0 vs. 0.6 points), while reducing upload by 30.6-49.0$\times$. In federated classification, DSL reaches 96.83% on MNIST with 8.9$\times$ less uplink than FedAvg and provides lower-communication operating points on CIFAR-10 and Tiny ImageNet. These results establish random-carrier outputs as a practical communication primitive for knowledge sharing across distinct collaborative learning paradigms.
CommentsAccepted to the Collaborative, Open, and DECentralized training of Foundation Models workshop at NeurIPS 2026