PulseAugur
实时 22:43:10
Deutsch(DE) NVIDIA-Nemotron-Labs-Teacher-STEM: 550B Parameter (55B aktiv), LatentMoE mit Mamba-2, MoE und Multi-Token Prediction, bis 1M Tokens Kontext. Laut Model Card auf

NVIDIA发布具有100万上下文的Nemotron-Labs-Teacher模型 · 跟踪4个来源

NVIDIA发布了一系列Nemotron-Labs-Teacher模型,每个模型拥有5500亿参数,但只有550亿参数被激活使用。这些模型采用了包含Mamba-2、MoE和多Token预测的LatentMoE架构,支持高达100万Token的上下文窗口。这些模型在OpenMDW-1.1许可下可用,并且需要大量的硬件支持,例如4块B200/GB200或8块H100 GPU才能运行。 AI

影响 这些模型在上下文窗口长度和架构方面推动了界限,可能影响未来的LLM开发。

排序理由 前沿实验室模型发布,附带系统卡。

在 Mastodon — mastodon.social 阅读 →

AI 生成摘要 · Google Gemini · 来自 4 个来源。 我们如何撰写摘要 →

NVIDIA发布具有100万上下文的Nemotron-Labs-Teacher模型 · 跟踪4个来源

报道来源 [4]

  1. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    NVIDIA-Nemotron-Labs-Teacher-STEM: 550B Parameters (55B active), LatentMoE with Mamba-2, MoE and Multi-Token Prediction, up to 1M Tokens Context. According to the Model Card on

    NVIDIA-Nemotron-Labs-Teacher-STEM: 550B Parameter (55B aktiv), LatentMoE mit Mamba-2, MoE und Multi-Token Prediction, bis 1M Tokens Kontext. Laut Model Card auf GPQA, MMLU-Pro und LiveCodeBench v6 auf Niveau von DeepSeek V4 Pro. Läuft ab 4xB200/GB200 oder 8xH100, Lizenz OpenMDW-1…

  2. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    NVIDIA Releases Nemotron-Labs-Teacher-General-Reasoning: 550B Parameters (55B Active), LatentMoE with Mamba-2, MoE and Attention plus Multi-Token-Prediction

    NVIDIA veröffentlicht Nemotron-Labs-Teacher-General-Reasoning: 550B Parameter (55B aktiv), LatentMoE mit Mamba-2, MoE und Attention plus Multi-Token-Prediction, Kontext bis 1M Tokens. Lizenz OpenMDW-1.1, Mindestanforderung 4x B200/GB200 oder 8x H100. https:// huggingface.co/nvidi…

  3. Mastodon — mastodon.social TIER_1 English(EN) · aisyndicate ·

    Nemotron-Labs-Teacher-Competition-Coding: 550B Parameter (55B aktiv), LatentMoE-Hybrid mit Mamba-2, MoE, Attention und Multi-Token Prediction. Coding-Teacher fü

    Nemotron-Labs-Teacher-Competition-Coding: 550B Parameter (55B aktiv), LatentMoE-Hybrid mit Mamba-2, MoE, Attention und Multi-Token Prediction. Coding-Teacher für Distillation, 1M Token Kontext, Lizenz OpenMDW-1.1, ab 8x H100. https:// huggingface.co/nvidia/NVIDIA-N emotron-Labs-T…

  4. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    NVIDIA-Nemotron-Labs-Teacher-Chat has 550B Parameters (55B Active), LatentMoE Architecture with Mamba-2, MoE, and Multi-Token Prediction. Context up to 1M Tokens, License

    NVIDIA-Nemotron-Labs-Teacher-Chat hat 550B Parameter (55B aktiv), LatentMoE-Architektur mit Mamba-2, MoE und Multi-Token Prediction. Kontext bis 1M Token, Lizenz OpenMDW-1.1, Mindesthardware 4x B200 oder 8x H100. https:// huggingface.co/nvidia/NVIDIA-N emotron-Labs-Teacher-Chat #…