A new study systematically evaluated 41 open-weight language models for zero-shot intent classification, assessing their performance across various constraints including compute, latency, and robustness. The research analyzed models ranging from 135M to 9B parameters across eight English datasets and one auxiliary five-shot dataset. Key findings indicate that instruction-tuned 3B models can outperform some 7B base models, differences among top models on the MASSIVE dataset are statistically indistinguishable, and popular benchmarks like SNIPS are saturated. The study also noted that instruction tuning's impact on confidence calibration is inconsistent. AI
IMPACT Provides practical guidance for selecting and evaluating open-weight language models for intent classification tasks.
RANK_REASON Academic paper presenting a systematic evaluation of multiple models on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]
- 41 Models
- ATIS
- Hugging Face
- Massive
- open-weight language models
- Parishruthi Ganesh
- Zero-Shot Intent Classification
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →