GPT-2 124M
PulseAugur coverage of GPT-2 124M — every cluster mentioning GPT-2 124M across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
Daedalus-150M: New hybrid LLM architecture optimized for CPU inference
Researchers have developed Daedalus-150M, a novel language model architecture optimized for CPU inference. Unlike traditional models that are scaled down after design, Daedalus-150M was built with CPU constraints in min…
-
New Daedalus-150M model achieves faster CPU inference with hybrid architecture
Researchers have developed Daedalus-150M, a novel language model optimized for efficient CPU inference. This hybrid model combines sparse attention with short convolutions, allowing two-thirds of its architecture to avo…
-
New ELO algorithm enhances learned optimizers for long-horizon tasks
Researchers have developed a new meta-training algorithm called ELO (Efficient Long-hOrizon) to improve learned optimizers (LOs). ELO addresses the challenges of scaling meta-training to long-horizon problems and compet…
-
Self-training restructures language models, research finds
A new research paper challenges the common understanding of self-training in language models, suggesting it restructures rather than flattens language. The study found that while surface-level linguistic features like d…