PulseAugur
EN
LIVE 08:22:28

Commercial LLMs outperform open-weight models across 30 languages, study finds

A new study published on arXiv investigates the performance of large language models (LLMs) across 30 languages, including the 24 official EU languages. The research found that commercial LLMs consistently outperform open-weight models, even in languages with significantly less available web text. The study also revealed that non-English languages are more expensive to run and score lower on average compared to English, despite commercial systems maintaining an advantage globally. AI

IMPACT Highlights the significant gap in multilingual capabilities between commercial and open-weight LLMs, suggesting current open models are insufficient for global language parity.

RANK_REASON The cluster contains an academic paper evaluating LLM performance on specific benchmarks. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Commercial LLMs outperform open-weight models across 30 languages, study finds

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Sherzod Hakimov, Karl Osswald, Jelle Psurek, Eszter Bukovszky, A. Altar L\"user, David Schlangen ·

    Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+

    arXiv:2608.01395v1 Announce Type: new Abstract: We evaluate large language models (LLMs) as language agents playing goal-directed dialogue games in self-play across 30 languages: the 24 official EU languages plus six others. Unlike static or preference-based evaluation, this para…