A new research paper investigates the phenomenon of "few-shot degradation" in language models, where providing examples can sometimes harm performance instead of improving it. The study, which tested 12 open-weight models on Ukrainian news classification and legal case outcome prediction tasks, found that the degradation effect is highly dependent on the specific task. Researchers developed a new metric called "content delta" to isolate the impact of demonstration content from prompt length, revealing that changes in model representations due to demonstration content, rather than prompt length, are key predictors of few-shot performance. AI
IMPACT This research offers a new way to understand and potentially mitigate performance issues when using few-shot prompting with language models.
RANK_REASON Research paper published on arXiv detailing findings about language model behavior. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Few-Shot Degradation
- Few-shot learning
- Language Models
- legal case outcome prediction
- Llama 3.3-70B
- News Classification from Social Media Using Twitter-based Doc2Vec Model and Automatic Query Expansion
- Ukrainian
- Volodymyr Ovcharov
- zero-shot learning
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →