nanoGPT
PulseAugur coverage of nanoGPT — every cluster mentioning nanoGPT across labs, papers, and developer communities, ranked by signal.
- 2026-05-15 research_milestone AI agents achieved new records in the nanoGPT training speedrun benchmark, surpassing human performance. source
7 day(s) with sentiment data
-
ScoutGPT uses language modeling for football player valuation
Researchers have developed ScoutGPT, a generative model that treats football match events as sequential tokens within a language modeling framework. This approach, utilizing a NanoGPT-based Transformer architecture, lea…
-
Sphere Retraction Normalizations generalize deep neural network training
Researchers have introduced Sphere Retraction Normalizations, a new framework for training deep neural networks that generalizes existing residual connection methods. This approach recasts residual connections on a Riem…
-
New SignMuon method compresses AI model updates to one bit per parameter
Researchers have developed SignMuon, a method for compressing model updates to a single bit per parameter, significantly reducing communication overhead. While SignMuon outperforms SignSGD in practice, it can still dive…
-
LLM API pricing comparison pitfalls: NanoGPT vs. aggregators
The article discusses the complexities of comparing pricing across different LLM API aggregators, using NanoGPT as an example. It highlights that seemingly transparent pricing can be misleading because different service…
-
METR metric quantifies AI agent cost-effectiveness vs. humans
METR has developed a new metric called the "expenditure horizon" to quantify the cost-effectiveness of AI agents. This metric aims to determine the point at which employing AI agents becomes more expensive than utilizin…
-
New spectral cap method enhances LLM training by controlling weight matrix geometry
Researchers have proposed a new method called an "Isotropy-Preserving Spectral Cap" to improve the training of large language models (LLMs). This technique aims to control the internal geometry of weight matrices during…
-
New metric quantifies AI optimization cost-effectiveness
Researchers have introduced a new metric called "expenditure horizon" to quantify an AI agent's optimization ability. This metric estimates the budget at which AI becomes more cost-effective than human effort for specif…
-
Users ditch ChatGPT for privacy-focused alternatives like Claude and NanoGPT
Two users describe their experiences switching away from ChatGPT for data analysis and general use, prioritizing privacy and cost savings. One user found Claude to be a better fit for data analysis, while another adopte…
-
117M Silia model trained in 5 hours on H100 GPU
A 117 million parameter Silia model was trained in just 5 hours on an H100 GPU, utilizing the synth-100M dataset. The model's architecture, detailed in a research paper, includes multi-headed attention and rotary positi…
-
New 'Radial Suppression' method accelerates neural network generalization
Researchers have developed a novel method called Radial Suppression to accelerate algorithmic generalization in neural networks. This technique addresses the common issue where models memorize training data before gener…
-
Gradient-free EntropyBeam model outperforms nanoGPT on Shakespeare benchmark
A new language model called EntropyBeam has demonstrated superior performance on the nanoGPT Shakespeare benchmark, achieving lower cross-entropy than the nanoGPT model. EntropyBeam operates without trainable parameters…
-
BeamGPT operator enhances language model training efficiency
A novel operator called BeamGPT has been developed, which significantly improves learning curves in language models by identifying sequence structures that standard attention mechanisms miss. This operator, when integra…
-
Developer implements GPTQ quantization from scratch, achieving minimal performance loss
A developer detailed their process of implementing the GPTQ quantization method from scratch on a nanoGPT model. This technique reduces model size and speeds up inference by lowering the precision of weights, but unlike…
-
Aurora optimizer enhances MLP training, outperforming Muon
Researchers have introduced Aurora, a novel spectral optimizer designed to address issues with non-uniform row norms in matrix parameters, particularly within MLP layers. This problem can lead to neurons receiving insuf…
-
New AngularMuown optimizer improves Transformer pre-training
Researchers have introduced AngularMuown, a novel optimization algorithm that implicitly performs angular step-size decay, building upon the principles of matrix-aware optimizers like Muon and Muown. This new method exp…
-
Hybrid LLM-GNN Model Enhances Quantum Circuit Optimization
A developer has created a hybrid model combining Large Language Models (LLMs) and Graph Neural Networks (GNNs) to improve the efficiency of the ADAPT-QAOA algorithm for optimizing quantum circuits. This approach aims to…
-
Student proposes Silia Transformer for parameter-efficient small models
A student researcher has introduced "Silia," a novel Transformer architecture designed for parameter efficiency in models under 10 million parameters. The architecture aims to combine the dynamic mixing of attention mec…
-
New 'Muon' optimization technique flattens matrix gradients
A new research paper introduces "Muon," an optimization technique that replaces matrix gradients with their polar factors. This method maintains singular directions but flattens the update spectrum, which the authors su…
-
Kronecker Embeddings slash language model parameters, boost performance
Researchers have developed Kronecker Embeddings, a novel method for representing tokens in language models that significantly reduces the number of trainable parameters. This approach replaces large embedding tables wit…
-
Community project proposed for training LLMs on 8GB VRAM consumer hardware
A user on r/LocalLLaMA is proposing a community project to train a large language model from scratch using only consumer-grade hardware, specifically targeting an 8GB VRAM limit. The goal is to create an accessible, fre…