Researchers have developed QK-Wanda, a novel method for pruning large language models that improves upon the existing Wanda technique. QK-Wanda couples the queries and keys within linear projections, leading to a significant reduction in reconstruction error compared to Wanda. While QK-Wanda is slightly slower than Wanda, it demonstrates substantial downstream performance gains on models like Llama 2 70B, improving perplexity and zero-shot accuracy. However, the effectiveness of QK-Wanda varies across different models, as seen with Llama 3.1 70B, indicating that local reconstruction error is not always a perfect predictor of overall model quality. AI
IMPACT Introduces a more effective pruning technique that could lead to more efficient deployment of large language models.
RANK_REASON The cluster contains a research paper detailing a new method for pruning large language models. [lever_c_demoted from research: ic=1 ai=1.0]
- A100
- H200
- Llama 2
- Llama 2 70B
- Llama 3
- Llama 3.1 70B
- QK-Wanda
- Qwen2.5
- Qwen2.5-72B
- TinyLlama
- Wanda
- WikiText-2
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →