PulseAugur
EN
LIVE 22:43:10
Deutsch(DE) NVIDIA-Nemotron-Labs-Teacher-STEM: 550B Parameter (55B aktiv), LatentMoE mit Mamba-2, MoE und Multi-Token Prediction, bis 1M Tokens Kontext. Laut Model Card auf

NVIDIA releases Nemotron-Labs-Teacher models with 1M context · 4 sources tracked

NVIDIA has released a suite of Nemotron-Labs-Teacher models, each with 550 billion parameters, though only 55 billion are actively used. These models leverage a LatentMoE architecture incorporating Mamba-2, MoE, and Multi-Token Prediction, supporting context windows up to 1 million tokens. The models are available under the OpenMDW-1.1 license and require significant hardware, such as 4x B200/GB200 or 8x H100 GPUs, for operation. AI

IMPACT These models push the boundaries of context window length and architecture, potentially influencing future LLM development.

RANK_REASON Frontier-lab model release with system card.

Read on Mastodon — mastodon.social →

AI-generated summary · Google Gemini · from 4 sources. How we write summaries →

NVIDIA releases Nemotron-Labs-Teacher models with 1M context · 4 sources tracked

COVERAGE [4]

  1. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    NVIDIA-Nemotron-Labs-Teacher-STEM: 550B Parameters (55B active), LatentMoE with Mamba-2, MoE and Multi-Token Prediction, up to 1M Tokens Context. According to the Model Card on

    NVIDIA-Nemotron-Labs-Teacher-STEM: 550B Parameter (55B aktiv), LatentMoE mit Mamba-2, MoE und Multi-Token Prediction, bis 1M Tokens Kontext. Laut Model Card auf GPQA, MMLU-Pro und LiveCodeBench v6 auf Niveau von DeepSeek V4 Pro. Läuft ab 4xB200/GB200 oder 8xH100, Lizenz OpenMDW-1…

  2. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    NVIDIA Releases Nemotron-Labs-Teacher-General-Reasoning: 550B Parameters (55B Active), LatentMoE with Mamba-2, MoE and Attention plus Multi-Token-Prediction

    NVIDIA veröffentlicht Nemotron-Labs-Teacher-General-Reasoning: 550B Parameter (55B aktiv), LatentMoE mit Mamba-2, MoE und Attention plus Multi-Token-Prediction, Kontext bis 1M Tokens. Lizenz OpenMDW-1.1, Mindestanforderung 4x B200/GB200 oder 8x H100. https:// huggingface.co/nvidi…

  3. Mastodon — mastodon.social TIER_1 English(EN) · aisyndicate ·

    Nemotron-Labs-Teacher-Competition-Coding: 550B Parameter (55B aktiv), LatentMoE-Hybrid mit Mamba-2, MoE, Attention und Multi-Token Prediction. Coding-Teacher fü

    Nemotron-Labs-Teacher-Competition-Coding: 550B Parameter (55B aktiv), LatentMoE-Hybrid mit Mamba-2, MoE, Attention und Multi-Token Prediction. Coding-Teacher für Distillation, 1M Token Kontext, Lizenz OpenMDW-1.1, ab 8x H100. https:// huggingface.co/nvidia/NVIDIA-N emotron-Labs-T…

  4. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    NVIDIA-Nemotron-Labs-Teacher-Chat has 550B Parameters (55B Active), LatentMoE Architecture with Mamba-2, MoE, and Multi-Token Prediction. Context up to 1M Tokens, License

    NVIDIA-Nemotron-Labs-Teacher-Chat hat 550B Parameter (55B aktiv), LatentMoE-Architektur mit Mamba-2, MoE und Multi-Token Prediction. Kontext bis 1M Token, Lizenz OpenMDW-1.1, Mindesthardware 4x B200 oder 8x H100. https:// huggingface.co/nvidia/NVIDIA-N emotron-Labs-Teacher-Chat #…