Two new research papers explore methods to improve the quantization of large language models (LLMs) while preserving their performance and uncertainty behavior. The first paper, "Jacobian-guided Noise Injection for Quantization Robustness in Large Language Models," proposes a training strategy that injects Gaussian noise into pre-attention logits, with variance determined by the Jacobian Frobenius norm, showing significant gains on image and text benchmarks. The second paper, "Target-Aware Calibration Data Selection for Preserving Uncertainty in Quantized Language Models," introduces Doubt-Preserving Quantization (DPQ), a pre-quantization technique that selects calibration data based on specific deployment targets to maintain uncertainty behavior across various NLP tasks. AI
IMPACT These papers offer new techniques to make LLMs more efficient for deployment by improving quantization methods, potentially leading to wider adoption and reduced computational costs.
RANK_REASON Two academic papers published on arXiv detailing novel methods for LLM quantization.
- arXiv
- Deepanshu Pandey
- Doubt-Preserving Quantization
- Hugging Face
- ImageNet-1K
- Jacobian-guided Noise Injection
- large-language models
- SQuAD2.0
- WikiText
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →