VIDRAFT has released POCKET-Darwin-180B, a 4-bit quantized version of their 180-billion-parameter Darwin-180B-RSI model. This version is compatible with llama.cpp and can run on consumer hardware, including laptops without dedicated GPUs, by utilizing sparse Mixture-of-Experts routing and a novel graft quantization technique. The model's size has been reduced from 360 GB to 111 GB, allowing for local inference with significantly lower hardware costs and system RAM requirements compared to traditional enterprise GPU setups, while reportedly maintaining benchmark accuracy. AI
IMPACT Enables local inference of large models on consumer hardware, reducing costs and increasing accessibility for developers.
RANK_REASON Frontier-lab model release with system card [lever_c_demoted from frontier_release: ic=1 ai=1.0]
- Darwin-180B-RSI
- GGUF
- llama.cpp
- mixture of experts
- MMLU-Pro
- NVIDIA H100
- POCKET-Darwin-180B
- Qwen3.8 Flash-Next
- VIDRAFT
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →