PulseAugur
EN
LIVE 07:57:44

Bonsai 27B models tested on Terminal-Bench 2.0, 1-bit version unusable

A user tested the Ternary-Bonsai-27B (2-bit) and Bonsai-27B (1-bit) models on the Terminal-Bench 2.0 benchmark, finding that the 2-bit version achieved a score of 7.9%. This performance was lower than the Qwen3.5-9B model, which also fits within 8GB of VRAM. The 1-bit Bonsai model proved unusable in an agentic harness due to non-termination issues, though it performed adequately on simple prompts. AI

IMPACT Demonstrates trade-offs between extreme quantization and model performance in agentic tasks.

RANK_REASON User benchmark of open-source models on a specific harness. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Bonsai 27B models tested on Terminal-Bench 2.0, 1-bit version unusable

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Creative-Regular6799 ·

    I ran Ternary-Bonsai-27B (2-bit) and Bonsai-27B (1-bit) on Terminal-Bench 2.0, in 8GB VRAM

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1v1ya97/i_ran_ternarybonsai27b_2bit_and_bonsai27b_1bit_on/"> <img alt="I ran Ternary-Bonsai-27B (2-bit) and Bonsai-27B (1-bit) on Terminal-Bench 2.0, in 8GB VRAM" src="https://preview.redd.it/315dccgwageh1.jpe…