PulseAugur
EN
LIVE 10:46:40

ClipProj reduces MiniMax H3 VRAM needs by 70% using smaller Qwen3-VL models

A new set of projection matrices, ClipProj, has been developed to enable smaller Qwen3-VL models to replace the larger Qwen3-VL-32B text encoder in the MiniMax H3 diffusion model. This significantly reduces VRAM requirements from 15.7 GB to 4.5 GB without altering the core diffusion model, VAEs, or sampler. The projection matrices are trained using ridge regression, with a residual network being the only trainable component, and are designed to work with any variant of their respective model sizes. AI

IMPACT Enables running advanced diffusion models on hardware with significantly less VRAM, broadening accessibility.

RANK_REASON This is a research release of new model components and techniques for diffusion models. [lever_c_demoted from research: ic=1 ai=1.0]

Read on Hugging Face Trending Models →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

ClipProj reduces MiniMax H3 VRAM needs by 70% using smaller Qwen3-VL models

COVERAGE [1]

  1. Hugging Face Trending Models TIER_1 (HR) · NicoLab28 ·

    NicoLab28/ClipProj-MiniMax-H3

    text-to-video · 0 downloads · 71 likes