AI 中文总结
提出基于可信执行环境的联邦学习系统,首次提供外部可验证的中心化差分隐私保证,提升隐私-效用权衡并已投入生产用于Gboard模型训练。
AI 中文摘要
联邦学习(FL)允许拥有私有数据的设备协作训练共享模型。我们提出了一种基于可信执行环境(TEEs)的下一代联邦学习系统,解决了早期系统面临的运营挑战,并首次提供外部可验证的中心化差分隐私(DP)保证,同时实现更好的隐私-效用权衡。在我们的系统中,设备上传的数据使用由TEE托管的密钥管理服务(KMS)管理的密钥进行加密。上传的数据在密码学上绑定到一项策略,该策略限制后续可在服务器端TEE中处理数据的Python程序集合。外部各方可以检查公开的透明日志,以观察这些策略允许的工作负载集合。我们的实验结果表明,新系统通过将收集到的数据以优化DP保证且不受设备可用性影响的调度方式整合到服务器端工作负载中,提高了设备覆盖率,并有利地移动了隐私-效用曲线。我们的新系统已投入生产,使得Android键盘(Gboard)的模型能够更快训练,并在更小且现在可外部验证的隐私预算下,相比使用先前系统训练的模型获得更好的准确性。
英文摘要
Federated Learning (FL) allows devices with private data to collaborate in training a shared model. We present a next-generation FL system based on Trusted Execution Environments (TEEs) that addresses operational challenges associated with earlier systems and provides externally verifiable central Differential Privacy (DP) guarantees for the first time while offering a better privacy-utility tradeoff. In our system, devices upload data encrypted with keys managed by a TEE-hosted Key Management Service (KMS). The uploaded data is cryptographically tied to a policy limiting the set of Python programs that may later process the data in server-side TEEs. External parties may inspect public transparency logs to observe the set of workloads allowed by these policies. Our experimental results show that the new system improves device coverage and favorably shifts privacy-utility curves by enabling collected data to be integrated into the server-side workload at a schedule that optimizes DP guarantees and is unaffected by device availability. Our new system has been productionized, enabling models for the Android Keyboard (Gboard) to be trained faster and achieve better accuracy under smaller, now externally verifiable privacy budgets in comparison to models trained using the prior system.