Two new research papers propose novel methods for post-training quantization (PTQ) of large language models, aiming to reduce their size and computational requirements. The first paper, "From Sweep to Seam: Interleaved Cross-Block Post-Training Quantization," introduces Interleaved Cross-Block Quantization (ICBQ), which refines quantization boundaries by revisiting them twice. The second paper, "ReQuant: Fixed-Grid Discrete Refinement for Post-Training Quantization," presents ReQuant, a plug-and-play refinement procedure that iteratively optimizes discrete weight assignments after an initial quantization step. Both methods aim to improve the performance of quantized models, especially at lower bit-widths, and can be integrated with existing PTQ pipelines. AI
IMPACT These techniques could significantly reduce the computational and memory footprint of LLMs, making them more accessible and efficient for deployment.
RANK_REASON Two academic papers published on arXiv proposing new methods for LLM quantization.
- GPTAQ
- large-language models
- Post-training quantization (PTQ)
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- GPTQ
- Hugging Face
- Interleaved Cross-Block Quantization
- ScienceCast
- Sequential CBQ
- Transformer++
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →