arXiv cs.LGOctober 1, 2026
Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?
Excerpt
arXiv:2609.40335v1 Announce Type: new Abstract: Differentially Private Stochastic Gradient Descent (DP-SGD) is a leading approach for privacy-preserving fine-tuning of large language models (LLMs). Many decoder-only LLMs employ weight tying between input and output embeddings, a design choice originally introduced for parameter efficiency and improved language modeling performance in the non-private setting. However, the impact of weight tying under differentially private training remains largel