A user tested the Ternary-Bonsai-27B (2-bit) and Bonsai-27B (1-bit) models on the Terminal-Bench 2.0 benchmark, finding that the 2-bit version achieved a score of 7.9%. This performance was lower than the Qwen3.5-9B model, which also fits within 8GB of VRAM. The 1-bit Bonsai model proved unusable in an agentic harness due to non-termination issues, though it performed adequately on simple prompts. AI
IMPACT Demonstrates trade-offs between extreme quantization and model performance in agentic tasks.
RANK_REASON User benchmark of open-source models on a specific harness. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →