GPT-2 Medium
PulseAugur coverage of GPT-2 Medium — every cluster mentioning GPT-2 Medium across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New methods tackle selective unlearning in language models · 2 sources tracked
Researchers are developing advanced techniques for selective unlearning in language models to remove specific data without degrading overall performance. One method, GRAPHSU, uses a graph-guided approach to expand delet…
-
New research probes MUON optimizer's convergence and proposes MALT extension
Two new research papers explore the MUON optimization algorithm, a method used in training large language models. The first paper introduces MALT, an extension of MUON that incorporates lightweight diagonal precondition…
-
New research tackles evaluation and architecture for masked diffusion language models
Two new research papers introduce novel evaluation protocols and architectures for masked diffusion language models (MDLMs). The first paper, "CaRE," proposes a compute-aware framework to standardize evaluations, reveal…
-
New fMRI decoding method shows language models obscure failures
Researchers have developed a new method for decoding continuous language from fMRI signals, improving upon existing encoding pipelines with expanded voxel selection and a more advanced language model. They also introduc…
-
New 7B Uniform Diffusion Language Model 'Sumi' Released, Alongside Diffusion Model Advancements
Researchers have introduced Sumi, a 7-billion parameter uniform diffusion language model (UDLM) pretrained from scratch on 1.5 trillion tokens. This open-source model demonstrates competitive performance against autoreg…
-
AutoCompress method isolates critical transformer layers for efficient compression
Researchers have developed AutoCompress, a novel method for compressing transformer models by isolating and preserving the critical first layer (Layer 0). This approach, termed Critical Layer Isolation (CLI), showed tha…