A new study published on arXiv explores the impact of weight tying in decoder-only Large Language Models (LLMs) when fine-tuned using Differentially Private Stochastic Gradient Descent (DP-SGD). The research found that untying the input and output embeddings consistently outperformed weight-tied models, leading to accuracy gains of up to 4.74% on benchmarks like SST-2 and QNLI. Furthermore, untied embeddings facilitate more memory-efficient DP-SGD training by enabling the use of ghost clipping, resulting in over 60% lower memory usage compared to weight-tied models. AI
IMPACT Untied embeddings offer a more efficient and effective approach for private LLM fine-tuning, potentially influencing future model designs.
RANK_REASON The cluster contains a research paper detailing findings on LLM architecture and training methods.
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →