SmolLM3 3B
PulseAugur coverage of SmolLM3 3B — every cluster mentioning SmolLM3 3B across labs, papers, and developer communities, ranked by signal.
- 2026-05-12 research_milestone SmolLM3 3B achieved a top score in the Works With Agents benchmark for agent coding. source
3 day(s) with sentiment data
-
New dataset AMPLE-Math probes value of privileged info in LLM self-distillation
Researchers have developed AMPLE-Math, a new dataset comprising over 5,000 mathematical problems, to investigate the impact of privileged information in on-policy self-distillation (OPSD) for language models. Their find…
-
Instella-MoE: New open-source MoE language model released
A new technical report introduces Instella-MoE, an open-source Mixture-of-Experts (MoE) language model with 16 billion total parameters. Trained on AMD Instinct GPUs, the model incorporates innovations like Gated Multi-…
-
New recipe trains AI models on consumer GPUs for under $7,000
Researchers have developed a cost-efficient pretraining recipe for language models, enabling training on consumer-grade hardware like RTX 5090 GPUs for under $7,000. This new method, demonstrated with the Puro-2B model …
-
Open-source recipe trains 2B LLMs on consumer GPUs for under $7K
Researchers have developed an open-source pretraining recipe that significantly reduces the cost of training large language models, making them accessible on consumer GPUs for under $7,000. Their Puro-2B model, trained …
-
webAI releases TwIL-LM formal logic models for local autoformalization
webAI has launched TwIL-LM, a family of two formal logic models available in 1.7B and 3B parameter sizes. These models are designed for autoformalization, translating English into first-order logic and verifying conclus…
-
New research tackles LLM alignment, safety, and optimization challenges
Researchers are exploring new methods to improve the alignment and reliability of large language models (LLMs). One study identifies a vulnerability in byte-pair encoding (BPE) tokenization that can be exploited to bypa…
-
Quantization and temperature effects on LLM safety analyzed
A new study investigated the combined effects of model quantization and sampling temperature on the safety alignment of large language models. Researchers found that standard quantization methods like INT4 and INT8 gene…
-
New Sparse Autoencoder Model Enhances LLM Feature Interpretability
Researchers have introduced the Sign-Aware Gated Sparse Autoencoder (SA-GSAE), a novel architecture designed to improve the interpretability of features extracted from Large Language Models. Unlike standard SAEs that en…
-
Tiny models outperform frontier AI in agent coding benchmark
A recent agent coding benchmark revealed that smaller, more efficient models are outperforming larger, frontier models. The SmolLM3 3B model, capable of running on a laptop, achieved a score of 93.3, significantly surpa…