Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?
Untied embeddings beat weight tying for decoder-only LLMs under DP-SGD, raising accuracy and cutting memory by over 60%.
Researchers study whether tying input and output embeddings still helps decoder-only LLMs when fine-tuning with DP-SGD. On GPT-2 and DistilGPT2, untied embeddings consistently beat tied models, with accuracy gains up to 4.74 percentage points on SST-2, QNLI, and QQP. Untying also allows memory-efficient ghost clipping, while tied weights complicate ghost-norm computation. Untied models used over 60% less memory while keeping ghost clipping's benefits.
- Untied embeddings gained up to 4.74 points on SST-2, QNLI, and QQP.
- Tests used GPT-2 and DistilGPT2 trained with DP-SGD.
- Untying enables ghost clipping and cuts memory by over 60%.
- Weight tying complicates ghost-norm computation and reduces its benefit.
Full article192 words · extracted from arxiv.org · click to collapse
Differentially Private Stochastic Gradient Descent (DP-SGD) is a leading approach for privacy-preserving fine-tuning of large language models (LLMs). Many decoder-only LLMs employ weight tying between input and output embeddings, a design choice originally introduced for parameter efficiency and improved language modeling performance in the non-private setting. However, the impact of weight tying under differentially private training remains largely unexplored. In this work, we investigate the role of weight tying in the DP setting using GPT2 and DistilGPT2 as representative decoder-only architectures. Interestingly, we find that untied embeddings consistently outperform weight-tied models under DP-SGD, achieving gains of up to 4.74% points in accuracy on SST-2, QNLI, and QQP. Beyond improved utility, untying embeddings enables the use of memory-efficient ghost clipping for DP-SGD. By contrast, weight tying introduces shared-parameter interactions that complicate standard ghost norm computation and largely negate its computational advantages. As a result, untied models achieve over 60% lower memory usage while preserving the benefits of ghost clipping. Our results indicate that untied embeddings provide a more effective and scalable design for differentially private training of decoder-only LLMs and highlight the need to revisit standard LLM architectural choices in the privacy-preserving setting.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.40335