A new paper evaluates programmatic tool calling (PTC) against traditional JSON tool calling for large language models. The study found that PTC, which exposes tools as typed Python stubs for models to invoke, matches or exceeds JSON tool calling performance on the BFCL v4 benchmark for most models tested. Notably, the GPT-5.6 family saw a 10.6% improvement with PTC over the JSON baseline, and PTC remained stable under context rotation conditions where the baseline degraded. AI
IMPACT Programmatic tool calling offers a more robust and capable method for LLMs to interact with external tools, potentially enhancing agentic behavior and performance on complex tasks.
RANK_REASON Academic paper evaluating a new method for LLM tool use. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →