OLMo 3 7B
PulseAugur coverage of OLMo 3 7B — every cluster mentioning OLMo 3 7B across labs, papers, and developer communities, ranked by signal.
- 2026-05-13 research_milestone A paper was published detailing the early formation and refinement of persona vectors during LLM pretraining. source
3 day(s) with sentiment data
-
NCP-ArchPreview LLM advances beyond token prediction to concept prediction
Researchers have introduced NCP-ArchPreview, a novel large language model that moves beyond traditional next-token prediction to incorporate next-concept prediction. This approach trains the model to predict discrete co…
-
Instella-MoE: New open-source MoE language model released
A new technical report introduces Instella-MoE, an open-source Mixture-of-Experts (MoE) language model with 16 billion total parameters. Trained on AMD Instinct GPUs, the model incorporates innovations like Gated Multi-…
-
New research explores advanced fine-tuning techniques for LLMs · 3 sources tracked
Three new research papers explore advanced techniques for supervised fine-tuning (SFT) of large language models. The first paper investigates optimal hyperparameters like learning rate and batch size across different mo…
-
New research reveals "inverted" steering vectors in LLMs
Researchers have identified an "inverted detection-control phenomenon" in steering vectors (SVs), a technique used to influence the output of large language models. This phenomenon occurs when highly discriminative SVs,…
-
AMD releases open Instella-MoE-16B LLM with 2.8B active parameters
AMD has released Instella-MoE-16B-A3B, an open-source Mixture-of-Experts language model. This model features 16 billion total parameters but only activates 2.8 billion per token, utilizing architectural innovations like…
-
AI value dynamics studied using population genetics tools
Researchers have developed a method to measure and forecast how AI model values change over time, particularly in self-training feedback loops. This project, conducted over five weeks, uses tools from population genetic…
-
OLMo 3 7B training reveals structured harmfulness directions
Researchers have analyzed the development of harmfulness representations within the OLMo 3 7B model during its training process. They identified distinct but related linear activation directions for various harmfulness …
-
Persona vectors in LLMs form early in pretraining, study finds
Researchers have identified that specific behavioral traits, like sycophancy, are represented by 'persona vectors' within large language models. These vectors form very early in the pretraining process, within the first…