A new GGUF format called Q8_CR has been developed, aiming to optimize diffusion models for specific GPU architectures like Ampere and Turing. This format leverages ComfyUI's INT8 kernel and aims to minimize the quality difference between INT8 and FP16 checkpoints. Q8_CR stores linear weights as pre-rotated INT8 with FP32 scales, while preserving precision for sensitive tensors, offering improved inference speed and reduced VRAM usage compared to standard Q8_0 and FP16 formats. AI
IMPACT Offers improved inference speed and reduced VRAM usage for specific diffusion models on compatible hardware.
RANK_REASON This is a technical improvement to a file format for AI models, not a release from a frontier lab.
- Turing
- Ampere
- ComfyUI
- Flux 2 Klein
- GGUF
- Ideogram
- INT8 ConvRot
- Krea 2
- Q8_CR
- RTX 20xx
- RTX 30xx
- Z-Image
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →