A new method allows a 180-billion-parameter model, POCKET-Darwin-180B-GGUF, to run on consumer hardware without a dedicated GPU. This is achieved through a combination of the model's sparse mixture-of-experts architecture, which only activates a fraction of its parameters per token, and a selective 4-bit quantization process that preserves the accuracy of crucial weights. The quantized model, available in GGUF format, requires significantly less storage and memory, enabling it to run on laptops with as little as 8GB of VRAM or even on mini PCs with sufficient RAM, while maintaining comparable performance on benchmarks like MMLU-Pro. AI
IMPACT Enables running large language models on consumer hardware, democratizing access and reducing reliance on cloud infrastructure.
RANK_REASON Technical post detailing a method for running a large model on consumer hardware, including accuracy metrics and reproduction steps. [lever_c_demoted from research: ic=1 ai=1.0]
- Darwin-180B-RSI
- GGUF
- Hugging Face
- llama.cpp
- MMLU-Pro
- modelscope
- POCKET-Darwin-180B-GGUF
- RTX 5060 Laptop
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →