Distributed and Private Textual Data Synthesis from Embeddings
Researchers propose a distributed differentially private text synthesis method combining DP summaries and secure protocols, removing the need for a trusted curator.
The paper presents a differential privacy and cryptography co-design for synthesizing textual training data without a trusted curator or tightly synchronized user participation. It releases a one-time DP summary in embedding space, identifying frequent semantic regions and their DP centroids to enable training-free offline text synthesis, with semantic support protection to avoid exposing rare user texts. A custom secure protocol enforces end-to-end DP guarantees over distributed user data. Across four benchmarks the approach achieves utility comparable to the state-of-the-art centralized DP synthesis method.
- DP-cryptography co-design removes the trusted-curator assumption for distributed text synthesis
- Releases one-time DP embedding centroids, enabling training-free offline synthesis
- Semantic support protection keeps released summaries away from rare user texts
- Matches state-of-the-art centralized DP synthesis utility on four benchmarks
Full article195 words · extracted from arxiv.org · click to collapse
We revisit differentially private (DP) text synthesis in the realistic setting of distributed users, where privacy concerns preclude a trusted curator with access to raw user texts. Existing DP text synthesis pipelines are designed for a trusted, centralized curator and often cannot be deployed in distributed settings due to unrealistic trust and access assumptions; when adapted naively, they require repeated, tightly synchronized user participation and incur significant overhead. To address this gap, we propose a DP--cryptography co-design for textual data synthesis that requires no trusted curator and requires only lightweight user participation. Our approach has two optimized components. First, we design a distributed-friendly DP synthesis algorithm that releases a one-time DP summary in an embedding space: it identifies frequent semantic regions and releases their DP centroids, enabling training-free, non-iterative offline text synthesis. We further introduce semantic support protection, which ensures the released summary avoids semantic neighborhoods of infrequent texts, reducing the risk of exposing rare user data. Second, we develop a custom secure protocol that implements this algorithm over distributed user data, enforcing end-to-end DP guarantees without requiring a trusted curator. On four benchmarks, we achieve utility comparable to the state-of-the-art centralized DP synthesis method.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.10104