Researchers have developed a new method called ternary multiplicative adaptation for fine-tuning transformers that are quantized to ternary weights. This approach uses a low-rank Kronecker factorization to represent discrete updates to ternary weights, allowing for parameter-efficient adaptation without dequantization. Experiments on models like ternarized LLaMA-3 and ViT-B/16 show that this method significantly improves performance compared to existing low-bit and ternary baselines. AI
IMPACT Enables more efficient fine-tuning of highly quantized models, potentially reducing computational costs for AI development.
RANK_REASON The cluster contains a research paper detailing a new method for fine-tuning quantized transformer models. [lever_c_demoted from research: ic=1 ai=1.0]
- Alexandru-Dragos Manolache
- arXiv
- Hugging Face
- LLaMA-3
- low-rank Kronecker factorization
- ternary multiplicative adaptation
- Ternary transformers
- transformers
- ViT-B/16
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →