A new version, v3, of the MiniMax H3 text-to-image model has been released, featuring smaller text encoders (4B or 8B parameters) that aim to reduce model size while maintaining quality. These smaller encoders are mapped to the original 32B encoder's output using projection matrices, allowing the core DiT model to remain unchanged. Performance metrics show the 8B encoder achieving a cosine similarity of 0.9449 and the 4B encoder reaching 0.9381 against the 32B model, with improvements also noted in vision token conditioning. AI
IMPACT Reduces model size and potentially inference costs for text-to-image generation.
RANK_REASON This is an update to an existing model, not a new frontier release, and focuses on technical implementation details.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →