NVIDIA has released a suite of Nemotron-Labs-Teacher models, each with 550 billion parameters, though only 55 billion are actively used. These models leverage a LatentMoE architecture incorporating Mamba-2, MoE, and Multi-Token Prediction, supporting context windows up to 1 million tokens. The models are available under the OpenMDW-1.1 license and require significant hardware, such as 4x B200/GB200 or 8x H100 GPUs, for operation. AI
IMPACT These models push the boundaries of context window length and architecture, potentially influencing future LLM development.
RANK_REASON Frontier-lab model release with system card.
Read on Mastodon — mastodon.social →
- GB200
- LatentMoE
- Mamba-2
- Nemotron-Labs-Teacher-Chat
- Nemotron-Labs-Teacher-Competition-Coding
- Nemotron-Labs-Teacher-General-Reasoning
- Nemotron-Labs-Teacher-STEM
- NVIDIA
- OpenMDW-1.1
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →