SparseGPT
PulseAugur coverage of SparseGPT — every cluster mentioning SparseGPT across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New pruning methods enhance LLM efficiency by preserving output differences
Researchers have introduced a new family of pruning methods called "difference-informed pruning" designed to improve the efficiency of large language models. These methods focus on preserving the differences between mod…
-
80B Qwen LLM runs on 4.3GB RAM via advanced compression
A breakthrough in running large language models on edge devices has been demonstrated with the Qwen 80B model, which was reportedly run using only 4.3GB of RAM. This achievement is attributed to advanced compression tec…
-
PALS method improves LLM pruning by adjusting layer sparsity
Researchers have developed PALS (Percentile-Aware Layerwise Sparsity), a novel method for pruning large language models. Unlike existing one-shot methods that apply uniform sparsity, PALS dynamically adjusts sparsity ra…