PulseAugur
EN
LIVE 22:08:59
ENTITY MXFP4

MXFP4

PulseAugur coverage of MXFP4 — every cluster mentioning MXFP4 across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
7
17 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
1
6 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

5 day(s) with sentiment data

RECENT · PAGE 1/1 · 17 TOTAL
  1. TOOL · CL_175285 ·

    DeepSeek V4 Flash quantized for DwarfStar inference engine

    A user has created and shared quantized versions of the DeepSeek V4 Flash model, specifically tailored for the DwarfStar (DS4) inference engine. These GGUF files aim to provide faster performance than standard llama.cpp…

  2. SIGNIFICANT · CL_173179 ·

    Moonshot AI releases Kimi K3, a 2.8T parameter open-weight MoE model

    Moonshot AI has released Kimi K3, a 2.8 trillion parameter open-weight Mixture of Experts (MoE) model. This model, featuring Kimi Delta Attention and other architectural innovations, offers improved scaling efficiency a…

  3. TOOL · CL_171587 ·

    Unsloth releases Kimi K3 GGUF models, including large MXFP4 version

    Unsloth has started releasing GGUF model files for Kimi K3. The initial releases include the MXFP4 model, which is 1.5 TB in size, and the mmproj component. These files are available for users to download and utilize.

  4. TOOL · CL_167429 ·

    MXAttention framework optimizes MXFP4 attention for video generation

    Researchers have developed MXAttention, a novel data-free post-training quantization framework designed to optimize MXFP4 attention in diffusion-based video generation models. This framework addresses numerical issues l…

  5. SIGNIFICANT · CL_166090 ·

    Moonshot AI releases open-weight Kimi K3 model, challenging Fable 5 and GPT-5.6 Sol

    Moonshot AI has released an open-weight version of its Kimi K3 model, utilizing the MXFP4 format and capable of processing 1.56TB of data. This release positions Kimi K3 as a competitor to advanced models like Fable 5 a…

  6. SIGNIFICANT · CL_166043 ·

    Moonshot releases Kimi K3, a 2.8T parameter multimodal model with 1M context

    Moonshot has released Kimi K3, a new 2.8 trillion parameter multimodal model featuring a 1 million token context window and native vision capabilities. The model demonstrates impressive speed, achieving 460 tokens per s…

  7. COMMENTARY · CL_150219 ·

    Kimi K3's massive scale demands intensive networking despite optimizations · 8 sources tracked

    SemiAnalysis reports that the Kimi K3 model, with its 2.8 trillion parameters, requires significant network bandwidth despite optimizations like Kimi Delta Linear Attention (KDA). The model's architecture necessitates t…

  8. TOOL · CL_129305 ·

    DynamiQ framework accelerates LLM training with optimized gradient synchronization

    Researchers have developed DynamiQ, a new framework designed to accelerate the training of large language models by optimizing gradient synchronization. This method addresses the network bottleneck issue in large-scale …

  9. SIGNIFICANT · CL_124570 ·

    GLM5.2 deployed on AMD MI355X for cheaper inference · 5 sources tracked

    Wafer.ai has successfully deployed GLM5.2 on AMD MI355X hardware, achieving a throughput of 2626 tokens/second/node and 213 tokens/second for single-stream inference. This deployment offers a cost advantage, with MI355X…

  10. MEME · CL_124528 ·

    User questions DeepSeek-V4 Flash quantization format

    A user on the r/LocalLLaMA subreddit is questioning the quantization format of the DeepSeek-V4 Flash model. The user points out that a Hugging Face repository by K. E. Bartowski lists the model as MXFP4, but the origina…

  11. TOOL · CL_118048 ·

    New W4A4 quantization technique enhances Wan2.2-I2V-A14B model inference

    Researchers have developed a novel W4A4 quantization technique for the Wan2.2-I2V-A14B model, aiming to improve inference efficiency on low-bit-width hardware. Their approach combines mixed precision for activation outl…

  12. RESEARCH · CL_79487 ·

    Paper catalogs 84 numeric formats for ML hardware consistency

    A new paper introduces a comprehensive catalog of 84 numeric formats used in machine learning hardware, addressing the challenge of silent divergences when porting models across different accelerators. The catalog inclu…

  13. TOOL · CL_41186 ·

    TORQ framework enhances LLM accuracy with MXFP4 quantization

    Researchers have developed TORQ, a new framework for quantizing Large Language Models (LLMs) using the MXFP4 format. This method addresses accuracy degradation issues by analyzing and correcting imbalances in activation…

  14. COMMENTARY · CL_37287 ·

    Benedict Evans analyzes AI's future; Huawei's HiFloat4 format shows promise

    Technology analyst Benedict Evans has released his 2026 analysis, exploring AI's transformative impact on business models and technology. He questions the current AI hype, discussing issues like 'enshittification' and t…

  15. RESEARCH · CL_37289 ·

    Huawei's HiFloat4 cuts AI training errors; CLI tools gain traction

    Huawei has developed a new 4-bit data format called HiFloat4, which reportedly reduces error rates by 33% compared to MXFP4 in AI model training on Ascend NPUs. This advancement is seen as a significant step in the tech…

  16. RESEARCH · CL_03577 ·

    llama.cpp and ik_llama.cpp add FP4 inference support for VRAM savings

    The llama.cpp and ik_llama.cpp projects have both integrated support for FP4 (4-bit floating-point) inference, a significant advancement for model quantization. llama.cpp now includes NVFP4, an Nvidia-specific format, w…

  17. RESEARCH · CL_00993 ·

    Huawei's HiFloat4 format boosts AI training efficiency; Anthropic automates safety research

    Huawei researchers have developed HiFloat4, a new 4-bit precision format for AI training and inference that outperforms existing formats like MXFP4 on Huawei's Ascend chips. This development is seen as a response to exp…