Researchers have developed AlphaMWE, a multilingual parallel corpus designed to aid research in machine translation and lexicography. This corpus includes annotations for multi-word expressions (MWEs) across six languages: Arabic, Chinese, English, German, Italian, and Polish. The creation process involved machine translation followed by human post-editing and annotation, with a focus on ensuring high quality through multiple review stages. A key finding from this work is that accurately translating MWEs remains a significant challenge for current state-of-the-art machine translation systems. AI
IMPACT This corpus could improve machine translation systems' ability to handle complex linguistic structures like multi-word expressions.
RANK_REASON This is a research paper describing the creation of a new dataset for NLP research. [lever_c_demoted from research: ic=1 ai=1.0]
- AlphaMWE
- Arabic
- Baidu Fanyi
- DeepL MT
- English
- German
- GoogleMT
- Italian
- Lifeng Han Dr
- machine translation
- Microsoft Bing Translator
- Polish
- Chinese
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →