A user tested 15 local large language models for their ability to use tools and agents, employing the Toolery benchmark. The benchmark involved 143 scenarios across four difficulty tiers, with each model undergoing 3 trials per scenario. Models were run locally using LM Studio with a 30k token context window and a temperature of 0.8. The qwen/qwen3.8-27b model achieved the highest overall score at 71.8%, excelling particularly in the Easy and Medium tiers. AI
IMPACT Provides insights into the practical capabilities of local LLMs for agentic tasks, guiding users on model selection for tool use.
RANK_REASON User-conducted benchmark and comparison of multiple open-source models. [lever_c_demoted from research: ic=1 ai=1.0]
- Bonsai 27B
- google/gemma-4-26b-a4b
- granite-4.2-8b
- lfm2.5-8b-a1b
- LM Studio
- mellum2-12b-a2.5b-thinking
- meta/muse-glimmer
- mistralai/devstral-small-2-2512
- ornith-1.5-35b-a3b
- ornith-1.5-9b
- qwen3.8-flash-next-iq2_xs
- qwen/qwen3.8-27b
- Toolery
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →