A recent benchmark tested two local large language models, Qwen3-14B and Llama-3.2-3B, on their ability to perform function calls for real-world API tasks. While both models demonstrated strong JSON validity, with the smaller Llama-3.2-3B model showing a slight edge, the larger Qwen3-14B model proved superior in argument correctness, especially in multi-turn conversations and error recovery scenarios. Despite falling short of GPT-4's reliability, the Qwen3-14B model's performance suggests it is better suited for complex local coding agents that require accurate data extraction and context maintenance. AI
IMPACT Local LLMs are improving in function calling, with larger models showing better accuracy in complex scenarios.
RANK_REASON Comparison of two specific LLMs on a technical capability (function calling). [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →