PulseAugur
EN
LIVE 15:16:18
ENTITY NVFP4

NVFP4

PulseAugur coverage of NVFP4 — every cluster mentioning NVFP4 across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
14
66 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
5
14 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

11 day(s) with sentiment data

RECENT · PAGE 1/5 · 100 TOTAL
  1. SIGNIFICANT · CL_260915 ·

    Nvidia releases GLM-5.3 with 1M context and MoE architecture

    Nvidia has released GLM-5.3, a new model utilizing a Mixture-of-Experts (MoE) architecture with 753 billion total parameters and 40 billion active parameters. This model features sparse attention mechanisms, enabling a …

  2. SIGNIFICANT · CL_246120 ·

    NVIDIA releases DeepSeek-V4-Pro-0813 with DSpark-Draft-Head

    NVIDIA has released DeepSeek-V4-Pro-0813 as an NVFP4 checkpoint, integrating a DSpark-Draft-Head. This release features MXFP4 experts from the draft module losslessly mapped to NVFP4, ensuring consistent quantization be…

  3. TOOL · CL_247090 ·

    Together AI optimizes ThunderKittens for NVIDIA Vera Rubin Blackwell GPUs

    Together AI has gained access to NVIDIA's Vera Rubin NVL72 platform, which is based on the Blackwell architecture. Their team has updated their ThunderKittens software to leverage new features of the Vera Rubin chip, sp…

  4. SIGNIFICANT · CL_242611 ·

    NVIDIA releases GLM-5.3-Flash and Qwen3.8-27B for Blackwell systems

    NVIDIA has released two new models, GLM-5.3-Flash and Qwen3.8-27B, optimized for their Blackwell systems. GLM-5.3-Flash, a 320B MoE model with 18B active parameters, supports multimodal tasks and a 1M context window, re…

  5. TOOL · CL_239382 ·

    Scale-QLoRA enables lossless merging of LLM adapters in 4-bit models

    A new research paper introduces Scale-QLoRA, a method for merging LoRA adapters into native 4-bit quantized LLMs without accuracy loss. Traditional merging methods can degrade performance, but Scale-QLoRA preserves the …

  6. TOOL · CL_237883 ·

    NInfer, llama.cpp, vLLM speed and quality compared for Qwen3.8-27B

    A user conducted a performance comparison of three inference engines—NInfer, llama.cpp, and vLLM—on a single RTX 5090 GPU using the Qwen3.8-27B model. The evaluation focused on quality and speed for a production content…

  7. SIGNIFICANT · CL_236662 ·

    Nvidia releases Qwen3.8-Flash-Next model optimized for Blackwell hardware

    Nvidia has released Qwen3.8-Flash-Next, a 125-billion parameter model, as an NVFP4 checkpoint. This model incorporates Hybrid Attention and Mixture-of-Experts (MoE) architectures. It is optimized to run via vLLM on Nvid…

  8. TOOL · CL_234496 ·

    New research tackles FP4 pretraining stability with 2D block scaling

    A new research paper introduces a method for stable FP4 pretraining by addressing a critical issue with transpose-invariant 2D block scaling. Previous methods using 1D scaling groups suffered from scale inconsistency wh…

  9. TOOL · CL_231345 ·

    New OCGQuant method enhances NVFP4 quantization for Llama 3 and Qwen 3

    Researchers have developed OCGQuant, a new post-training quantization method designed to improve the accuracy of NVFP4 (an efficient microscaling format for low-bit inference) by addressing issues with activation outlie…

  10. TOOL · CL_228715 ·

    A.X K2 language model debuts with 688B parameters and agentic focus

    A new technical report introduces A.X K2, a 688 billion parameter Mixture-of-Experts (MoE) language model designed for agentic applications. Despite being trained on fewer tokens than its predecessor, A.X K1, A.X K2 dem…

  11. TOOL · CL_227612 ·

    NVIDIA RTX PRO 6000 BSE vs H100 NVL: NVFP4 performance gains detailed

    A comparison between NVIDIA's RTX PRO 6000 Blackwell Server Edition (BSE) and the H100 NVL highlights the performance gains offered by the new NVFP4 four-bit weight format with two-level scaling. The RTX PRO 6000 BSE, f…

  12. TOOL · CL_227089 ·

    New DAMP technique slashes LLM memory use and boosts speed

    Researchers have developed a novel quantization technique called DAMP (Decay-Aware Mixed-Precision Recurrent-State Quantization) to reduce the memory footprint and improve the speed of large language models that use rec…

  13. SIGNIFICANT · CL_222413 ·

    NVIDIA releases quantized DeepSeek and Qwen LLMs for Blackwell hardware

    NVIDIA has released quantized versions of two large language models, DeepSeek-V4-Pro-0813 and Qwen3.8-2.4T-A95B, utilizing their NVFP4 quantization method. The DeepSeek model, with 1.65 trillion parameters, employs Hybr…

  14. TOOL · CL_220188 ·

    Nvidia touts DSX MaxLPS for maximizing AI compute density within fixed power budgets

    Nvidia is promoting its DSX MaxLPS power management technology, designed to optimize compute density within fixed data center power budgets. At Hot Chips 2026, the company highlighted how this system, when paired with i…

  15. TOOL · CL_218697 ·

    SGLang bug causes endless repetition in FP8 lm_head models

    A bug in SGLang versions prior to commit 5375babb causes endless repetition and empty responses when serving models with FP8 lm_head configurations, such as unsloth/Qwen3.8-27B-NVFP4. This issue arises because SGLang in…

  16. TOOL · CL_219457 ·

    Qwen3.8-27B model achieves near-BF16 performance with aggressive NVFP4 quantization

    A new fully quantized version of the Qwen3.8-27B model, named Qwen3.8-27B-QUASAR-NVFP4, has been released. This model utilizes a novel quantization-aware distillation (QAD) algorithm called QUASAR, achieving an aggressi…

  17. TOOL · CL_217546 ·

    MiniMax AI expands H3 ecosystem with new integration index

    MiniMax AI is expanding its H3 ecosystem with a new index called "Awesome MiniMax H3 Integrations." This resource tracks community-built projects leveraging the H3 model, ranging from local ComfyUI setups requiring 24GB…

  18. TOOL · CL_214167 ·

    Uniform INT4 quantization outperforms NVFP4 on real gradient tensors

    A study comparing quantization palettes for gradient tensors found that a uniform INT4 palette outperformed NVIDIA's NVFP4 palette on real-world training data. The research suggests that the random Hadamard rotation, of…

  19. TOOL · CL_213438 ·

    User quantizes LTX-2.5 Gemma-4 12B text encoder for ComfyUI

    A user has developed a quantized version of the LTX-2.5 Gemma-4 12B text encoder, specifically optimized for ComfyUI. This new NVFP4 format reduces VRAM usage and functions as a direct replacement for the original encod…

  20. TOOL · CL_212876 ·

    Error feedback harms Adam optimizer performance, study finds

    A recent study has revealed that error feedback, a technique used to mitigate precision loss in machine learning, performs poorly when combined with the Adam optimizer. While effective with Stochastic Gradient Descent (…