发表机构
School of Computer Science and Engineering; The Hebrew University of Jerusalem(计算机科学与工程学院; 耶路撒冷希伯来大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对低预算跨筒仓联邦主动学习,发现同分布数据查询选择更具挑战性,提出基于联邦表示学习的新框架,性能优于现有方法。
AI 中文摘要
联邦主动学习(FAL)解决数据隐私和标签稀缺的双重挑战,而缺乏全局数据视图会给协调查询选择带来额外障碍。我们研究低预算场景下的跨筒仓FAL,该场景中标注决策最为关键。我们从理论和实证两方面刻画了异质性反转现象:在低预算设置中,同分布(IID)数据需要更强的协调以避免冗余查询,而异质数据自然会促进多样性;当预算提高时,这一趋势会反转。因此,与标准联邦学习(FL)中异质性是主要挑战的观点相反,我们表明在FAL的查询选择中,同分布(IID)设置更具挑战性。基于这些发现,我们提出一种新的FAL框架,该框架利用联邦表示学习在共享嵌入空间中对齐客户端数据。这使得服务器可以对可选混淆的客户端嵌入执行全局协调的主动选择,而标注仍保留在每个客户端本地。尽管我们的框架在更具挑战性的低预算场景中运行,但它的性能超过了现有FAL方法,即使这些方法被给予大得多的标注预算,证明了隐私约束下集中协调的价值。
英文摘要
Federated Active Learning (FAL) addresses the dual challenges of data privacy and label scarcity, where the absence of a global data view introduces additional hurdles for coordinated query selection. We study cross-silo FAL in the low-budget regime, where annotation decisions are most critical. We characterize, both theoretically and empirically, a heterogeneity reversal: in low-budget settings, homogeneous (IID) data requires stronger coordination to avoid redundant queries, whereas heterogeneous data naturally promotes diversity; this trend reverses at higher budgets. Thus, in contrast to the standard federated learning (FL) narrative where heterogeneity is a primary challenge, we show that IID settings are more challenging for query selection in FAL. Motivated by these findings, we propose a new FAL framework that utilizes federated representation learning to align client data in a shared embedding space. This enables the server to perform globally coordinated active selection over optionally obfuscated client embeddings, while annotation remains local to each client. Although our framework operates in the more challenging low-budget regime, it achieves performance that surpasses existing FAL methods even when they are given substantially larger annotation budgets, demonstrating the value of centralized coordination under privacy constraints.
Comments16 pages