Researchers from Google and affiliated universities have developed ToolGrad, a novel framework for generating data to train large language models in tool usage. Unlike previous query-first methods, ToolGrad constructs a verified tool-use chain first and then generates a matching user query, significantly improving efficiency and accuracy. This approach achieved a 99.8% pass rate on data generation and, when used to fine-tune Gemma-3 models, demonstrated performance competitive with leading proprietary models on the Berkeley Function Calling Leaderboard. AI
IMPACT This framework could significantly accelerate the development and deployment of LLMs capable of reliably using external tools.
RANK_REASON The cluster describes a new research framework and dataset for LLM tool-use data generation, including performance benchmarks. [lever_c_demoted from research: ic=1 ai=1.0]
- Claude 4.5 Opus
- Gemini 2.5 Flash-Lite
- Gemma~3
- GPT-5
- Hugging Face
- RIKEN AIP
- Tohoku University
- ToolACE
- ToolBench
- ToolGrad
- University of Tokyo
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →