A quantized version of the Darwin-180B-RSI model, named POCKET-Darwin-180B-GGUF, has been released, allowing it to run on consumer hardware without a dedicated GPU. This 111 GB model utilizes a Mixture of Experts (MoE) architecture, meaning only a fraction of its 180 billion parameters are active for each token generation. The model can achieve up to 21 tokens per second on a CPU with sufficient RAM, or slower speeds on laptops with less RAM by leveraging SSDs for weight loading, while maintaining the original model's precision. AI
IMPACT Enables running large language models on consumer hardware, reducing reliance on cloud GPUs.
RANK_REASON Release of a quantized, open-source model for local inference.
- Darwin-180B-RSI
- FINAL-Bench
- GGUF
- Hugging Face
- llama.cpp
- MMLU-Pro
- POCKET-Darwin-180B-GGUF
- RTX 5060
- VIDRAFT
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →