A new study published on arXiv investigates the effectiveness of generative versus encoder-based models for Named Entity Recognition (NER) across eleven Indic languages. The research, conducted on the Naamapadam benchmark, found that encoder-based models like mBERT and XLM-R significantly outperformed generative architectures, including fine-tuned LLMs such as Gemma-2-2B. The study identified distinct language clusters based on model performance and offered deployment guidelines for low-resource NLP. AI
IMPACT Identifies limitations of current generative LLMs for low-resource languages and highlights the continued strength of encoder models for specific NLP tasks.
RANK_REASON Research paper published on arXiv detailing empirical study of NLP models. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Gemma 2-2B
- Hugging Face
- Indic languages
- multilingual-BERT
- Named Entity Recognition
- XLM-RoBERTa
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →