Researchers have developed a new method called ternary multiplicative adaptation for fine-tuning transformers that are quantized to ternary weights. This approach uses a low-rank Kronecker factorization to represent discrete updates to ternary weights, allowing for parameter-efficient adaptation without dequantization. Experiments on models like ternarized LLaMA-3 and ViT-B/16 show that this method significantly improves performance compared to existing low-bit and ternary baselines. AI
影响 Enables more efficient fine-tuning of highly quantized models, potentially reducing computational costs for AI development.
排序理由 The cluster contains a research paper detailing a new method for fine-tuning quantized transformer models. [lever_c_demoted from research: ic=1 ai=1.0]
- Alexandru-Dragos Manolache
- arXiv
- Hugging Face
- LLaMA-3
- low-rank Kronecker factorization
- ternary multiplicative adaptation
- Ternary transformers
- transformers
- ViT-B/16
AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →