NVIDIA Model Optimizer
PulseAugur coverage of NVIDIA Model Optimizer — every cluster mentioning NVIDIA Model Optimizer across labs, papers, and developer communities, ranked by signal.
-
Nvidia releases GLM-5.3-Flash NVFP4 multimodal model with 1M context
Nvidia has released the GLM-5.3-Flash NVFP4 model, a quantized version of ZAI's GLM-5.3-Flash. This multimodal Mixture-of-Experts model is designed for reasoning, coding, and agentic tasks, featuring a hybrid attention …
-
Unsloth accelerates Qwen3.8-Flash-Next and GLM-5.3-Flash performance
Unsloth has released updates that significantly accelerate the performance of Qwen3.8-Flash-Next and GLM-5.3-Flash models, offering up to 2x faster generation speeds and reduced token consumption. These improvements are…
-
NVIDIA releases Qwen-Image-Flash text-to-image model
NVIDIA has released Qwen-Image-Flash, a new text-to-image generation model. This model is a distilled version of Qwen/Qwen-Image, utilizing a four-step process and DMD2 distillation techniques from NVIDIA's FastGen, Mod…
-
NVIDIA releases Qwen-Image-Flash for fast text-to-image generation
NVIDIA has released the Qwen-Image-Flash model, a distilled version of the Qwen/Qwen-Image model designed for rapid text-to-image generation. This model utilizes a four-step DMD2 distillation process and is optimized fo…
-
NVIDIA unveils Kimi-K2.6-DFlash for Moonshot AI latency optimization
NVIDIA has introduced Kimi-K2.6-DFlash, a specialized draft head designed for Moonshot AI's Kimi-K2.6 model. This new component is optimized for speculative decoding using the NVIDIA Model Optimizer and is intended to r…