W4A4
PulseAugur coverage of W4A4 — every cluster mentioning W4A4 across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
LLM Quantization: More Than Just Bit Reduction
Quantization in large language models is a complex process involving more than just reducing bit precision. It encompasses four key decisions: notation, format, evaluation, and the resulting capacity gains. Different qu…
-
New QUASAR methods enhance LLM accuracy in low-bit quantization
Two new research papers introduce QUASAR, a novel method for improving the accuracy of quantized large language models. The first paper focuses on a training-free post-training quantization approach that addresses issue…
-
New QUADS technique stabilizes NVFP4 RL for MoE LLMs
Researchers have developed a new technique called QUADS to stabilize reinforcement learning (RL) for Mixture-of-Experts (MoE) Large Language Models using the NVFP4 low-precision format. They identified activation error,…
-
New COD-TDQ method boosts quantized Transformer performance for object detection
Researchers have developed a new method called COD-TDQ to improve the performance of Transformer-based models for camouflaged object detection (COD) when using aggressive post-training W4A4 quantization. They identified…
-
RTX Pro 4500 GPU tested with new PrismaQuant LLM quantization methods
A user on Reddit's r/LocalLLaMA subreddit shared their experience testing new quantization methods for large language models on an RTX Pro 4500 GPU. They encountered issues with a specific Sakamakismile model, experienc…
-
Int4 w4a4 quantization praised for minimal quality loss
A Reddit user shared their experience with Int4 w4a4, a quantization method for AI models, expressing that it is "insane." They presented two images, one generated with Int4 and another with Int8, noting that the visual…
-
New GoodQ method uses generative models for zero-shot object detector quantization
Researchers have developed GoodQ, a new pipeline for Zero-Shot Quantization-Aware Training (ZSQ-OD) that leverages off-the-shelf generative models to create training datasets. This method addresses challenges such as de…