Wanda
PulseAugur coverage of Wanda — every cluster mentioning Wanda across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New pruning methods enhance LLM efficiency by preserving output differences
Researchers have introduced a new family of pruning methods called "difference-informed pruning" designed to improve the efficiency of large language models. These methods focus on preserving the differences between mod…
-
New Super-Tuning method enhances LLM fine-tuning efficiency
Researchers have developed a new method called Super-Tuning, which aims to make fine-tuning large language models (LLMs) more efficient. This technique reuses saliency signals from model pruning to identify which parame…
-
PALS method improves LLM pruning by adjusting layer sparsity
Researchers have developed PALS (Percentile-Aware Layerwise Sparsity), a novel method for pruning large language models. Unlike existing one-shot methods that apply uniform sparsity, PALS dynamically adjusts sparsity ra…
-
New pruning method preserves LLM reasoning performance
Researchers have developed a new training-free method called Causal Attribution Pruning (CAP) to reduce the size of large language models while preserving their reasoning capabilities. CAP identifies and prunes less cri…
-
New research quantifies error propagation in compressed transformers
Researchers have developed a method to better understand and manage error propagation in compressed transformer models. By measuring the ratio of output to input error (rho) at each layer, they found that errors accumul…