The MiniMax H3 text encoder, quantized to NVFP4, has been released and is significantly smaller than its original 26.4 GB size, now fitting onto a single 16 GB graphics card. This model utilizes Qwen3-VL-32B as its text encoder and features a split transformer architecture. Discussions suggest that while smaller quantized versions are common, they may impact prompt accuracy, with FP8-mixed versions offering a better balance. AI
IMPACT This release offers a more accessible version of the MiniMax H3 text encoder, potentially enabling wider experimentation and use on consumer hardware.
RANK_REASON Release of a quantized model and discussion of its technical specifications and performance.
Read on Mastodon — mastodon.social →
- MiniMax H3
- Qwen3-VL-32B
- Int8
- Qwen3-VL-32B-Instruct-layer50_bf16.safetensors
- Administrador de Infraestructuras Ferroviarias
- Mastodon
- NVFP4
- Tono_Ken3
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →