Toward verifiably private learning from federated data
A TEE-based federated learning system trains Gboard models faster under smaller, verifiable privacy budgets.
The paper describes a TEE-based federated learning system in which devices upload data encrypted under keys from a TEE-hosted key management service. Uploads are cryptographically bound to policies that limit which Python programs may process them in server-side TEEs, and outsiders can inspect transparency logs of allowed workloads. The authors report better device coverage and privacy-utility tradeoffs because data can be incorporated on a schedule that optimizes differential privacy rather than device availability. The system is productionized for Android Keyboard (Gboard), training models faster with better accuracy under smaller, externally verifiable privacy budgets than the prior system.
- TEE-hosted KMS encrypts uploads and binds them to allowed Python workloads.
- Public transparency logs show which server-side TEE programs policies permit.
- Scheduling use independently of device availability improves the privacy-utility curve.
- Productionized for Gboard with faster training and smaller verifiable privacy budgets.
Full article189 words · extracted from arxiv.org · click to collapse
Federated Learning (FL) allows devices with private data to collaborate in training a shared model. We present a next-generation FL system based on Trusted Execution Environments (TEEs) that addresses operational challenges associated with earlier systems and provides externally verifiable central Differential Privacy (DP) guarantees for the first time while offering a better privacy-utility tradeoff. In our system, devices upload data encrypted with keys managed by a TEE-hosted Key Management Service (KMS). The uploaded data is cryptographically tied to a policy limiting the set of Python programs that may later process the data in server-side TEEs. External parties may inspect public transparency logs to observe the set of workloads allowed by these policies. Our experimental results show that the new system improves device coverage and favorably shifts privacy-utility curves by enabling collected data to be integrated into the server-side workload at a schedule that optimizes DP guarantees and is unaffected by device availability. Our new system has been productionized, enabling models for the Android Keyboard (Gboard) to be trained faster and achieve better accuracy under smaller, now externally verifiable privacy budgets in comparison to models trained using the prior system.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.31494