PulseAugur
EN
LIVE 23:57:47

New GGUF Q8_CR format optimizes diffusion models for specific GPUs

A new GGUF format called Q8_CR has been developed, aiming to optimize diffusion models for specific GPU architectures like Ampere and Turing. This format leverages ComfyUI's INT8 kernel and aims to minimize the quality difference between INT8 and FP16 checkpoints. Q8_CR stores linear weights as pre-rotated INT8 with FP32 scales, while preserving precision for sensitive tensors, offering improved inference speed and reduced VRAM usage compared to standard Q8_0 and FP16 formats. AI

IMPACT Offers improved inference speed and reduced VRAM usage for specific diffusion models on compatible hardware.

RANK_REASON This is a technical improvement to a file format for AI models, not a release from a frontier lab.

Read on r/StableDiffusion →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New GGUF Q8_CR format optimizes diffusion models for specific GPUs

COVERAGE [1]

  1. r/StableDiffusion TIER_2 English(EN) · /u/molbal ·

    Introducing GGUF Q8_CR - Mixing the best of GGUFs and INT8 ConvRot.

    <table> <tr><td> <a href="https://www.reddit.com/r/StableDiffusion/comments/1v7i5uj/introducing_gguf_q8_cr_mixing_the_best_of_ggufs/"> <img alt="Introducing GGUF Q8_CR - Mixing the best of GGUFs and INT8 ConvRot." src="https://preview.redd.it/em4pv8mhanfh1.png?width=140&amp;heigh…