Researchers have developed the AlphaMWE corpus to test the capabilities of large language models (LLMs) in machine translation, specifically focusing on Multiword Expressions (MWEs). The study evaluated 31 MT systems across various language pairs, including English to Chinese, Polish, German, and several Arabic dialects. Automatic evaluations using metrics like BLEU and BERT-score, followed by human evaluations, revealed that figurative language and MWEs continue to pose challenges for LLMs, and aggregate scores can mask language-specific errors. AI
IMPACT Highlights ongoing challenges for LLMs in nuanced language translation, suggesting areas for future model development.
RANK_REASON The item is a research paper detailing a new corpus and evaluation of LLM performance on machine translation tasks. [lever_c_demoted from research: ic=1 ai=1.0]
- AlphaMWE
- Arabic
- bert-score
- Bleu
- Egyptian Arabic
- English
- German
- LLMs
- machine translation
- Modern Standard Arabic
- Multiword Expressions
- Polish
- Standard Chinese
- Tunisian Arabic
- WMT2026
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →