发表机构
Coordinated Science Laboratory; University of Illinois Urbana-Champaign; University of Michigan(协调科学实验室; 伊利诺伊大学厄巴纳-香槟分校; 密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对跨孤岛联邦新站点加入时的类别扩展问题,提出三阶段自主流程,利用原型或本地记录证据,在有限无标注数据下保持新旧类别准确率平衡并优于基线。
AI 中文摘要
一个组织通常拥有的标注数据太少,无法训练出能够泛化的模型,而能够补充这些数据的记录却掌握在无法公开数据的其他组织手中。跨孤岛联邦学习提供了一条解决途径,因为参与者交换的是模型参数而非记录,但该方法通常需要预先确定两个方面的安排:参与站点以及模型能够预测的类别,而实际部署可能会打破这两点。在训练完成后,当原有站点已结束协作并离线时,一个新站点加入,其记录以无标注形式到达,混合了模型已识别的条件与所有参与者均未观察到的条件。我们提出了一种自主的三阶段流程,在加入站点上完全扩展模型:重构专家筛选新颖性,聚类将标记记录划分为候选条件,类均值描述旧类别,所有这些都在一个共享表示内完成。这些旧类别是从从未离开其所有者的记录中学习到的,因此常用的防遗忘手段不可用,而该流程提供了它们本应从两个不同来源之一携带的证据:联邦持有的原型或加入站点持有的记录。在一个真实的工业状态监测数据集上,端到端运行且不参考任何标签,任一来源都能将旧类别准确率保持在0.868或以上,遗忘率最多为0.063,两者差异为0.021,因此可以根据允许的披露程度而非实现的准确率来选择配置。两者都保持了旧类别与新类别准确率的平衡,而我们测量的所有其他替代方案都以牺牲一方为代价换取另一方,并且两者都比基于蒸馏和正则化的基线保留了更多旧类别。即使每个到达条件仅有6条标注记录且每个旧类别保留3条记录,这种平衡仍然成立。
英文摘要
An organization often holds too little labeled data to train a model that generalizes, and the records that would supply the rest sit with organizations that cannot release them. Cross-silo federated learning offers a way through, since participants exchange model parameters rather than records, but it ordinarily settles two aspects of the arrangement in advance, the participating sites and the classes the model can predict, and deployment can breach both. A new site joins after training, once the established sites have finished their engagement and gone offline, and its records arrive unlabeled, mixing conditions the model already recognizes with conditions no participant has observed. We present an autonomous three-stage procedure that expands the model entirely at the joining site: reconstruction experts screen for novelty, clustering separates the flagged records into candidate conditions, and class means describe the old classes, all inside one shared representation. Those classes were learned from records that never leave their owners, so the usual defenses against forgetting are unavailable, and the procedure supplies the evidence they would have carried from either of two dissimilar sources, prototypes held by the federation or records held by the joining site. On a real industrial condition-monitoring dataset, run end to end with no label consulted, either source holds old-class accuracy at 0.868 or above with forgetting at most 0.063, and the two differ by 0.021, so a configuration can be chosen by the disclosure it permits rather than the accuracy it delivers. Both keep old- and new-class accuracy in balance where every alternative we measure gives up one for the other, and both retain more of the old classes than distillation- and regularization-based baselines. The balance still holds with only 6 labeled records per arriving condition and 3 retained per old class.
CommentsAccepted to the Continual Learning for Enterprise AI Agents (CLEA) workshop at NeurIPS 2026. 9 pages, 1 figure, plus appendix