PrismML has released Bonsai 27B, a highly compressed version of Qwen3.6-27B, available in 1-bit and ternary variants. These models are designed to run on consumer hardware like laptops and phones, with the 1-bit version requiring only 3.9GB and the ternary version 5.9GB. Despite aggressive compression, the ternary model retains approximately 94.6% of the original FP16 model's performance, while the 1-bit version maintains 89.5%, making them suitable for applications requiring large context windows and efficient memory usage. AI
IMPACT Enables running large language models on consumer devices by drastically reducing memory footprint, potentially accelerating local AI deployment.
RANK_REASON Model release from a lab (PrismML) with specific performance metrics and hardware targets. [lever_c_demoted from frontier_release: ic=2 ai=1.0]
- AIME26
- Apache 2.0
- Apple Silicon
- BitNet
- Bonsai 27B
- DSpark
- EvalScope
- Gemma-4-31B
- FP16
- HQQ
- iOS
- LiveCodeBench
- MATH-500
- MMLU-Redux
- OpenAI
- PrismML
- Qwen3.6-27B
- vLLM
- Bonsai-27B-gguf
- llama.cpp
- Qwen 3.6 27b
- Ternary-Bonsai-27B-gguf
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →