A new arXiv paper investigates the performance disparity between English and Bengali in open large language models (LLMs). Researchers developed a consistent pipeline to translate 8 English benchmarks into Bengali and evaluated 10 open LLMs. The study found that script fragmentation, where tokenizers break down Bengali's alphasyllabary into smaller pieces, and evaluation format adherence contribute to the performance gap. While Bengali text costs significantly more tokens than English, this cost and sequence length showed only a weak correlation with model scores. AI
IMPACT Highlights potential biases in LLM evaluations and the challenges of multilingual model performance.
RANK_REASON Academic paper on LLM performance evaluation. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →