Researchers have developed multilingual deep learning models capable of identifying various text types, or "registers," on the open web across 16 languages. These models utilize the newly introduced Multilingual CORE corpora, comprising over 72,000 documents classified into 25 registers. The best-performing model achieved an F1 score of 79% on average across languages, demonstrating effectiveness even with a complex classification scheme. Performance significantly improved to over 90% F1 when documents with ambiguous labels were removed, indicating that inherent ambiguity in web registers, rather than model limitations, poses a challenge. Multilingual models consistently outperformed monolingual ones, especially for languages with less training data, though zero-shot performance on unseen languages showed a slight decrease. AI
IMPACT This research could improve the organization and understanding of vast amounts of web text, potentially aiding in content moderation, information retrieval, and cross-lingual analysis.
RANK_REASON The cluster contains an academic paper detailing new deep learning models and corpora for a specific NLP task. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →