Researchers have developed IHUBERT, a new Persian language model built on the RoBERTa-base encoder. This model was trained on a 45 GB curated dataset from the Sepahr-Danesh collection, totaling approximately 7-8 billion tokens. IHUBERT utilizes a multi-stage preprocessing pipeline, including semantic deduplication, to enhance corpus quality and balance domain representation. The model demonstrates strong performance across various Natural Language Understanding benchmarks, particularly excelling in extractive question answering tasks. AI
IMPACT Advances Persian language modeling capabilities and sets new benchmarks for NLU tasks in the Persian language.
RANK_REASON The cluster describes a new research paper detailing the creation and evaluation of a language model.
- DigiMag
- FarsTail
- Hugging Face
- IHUBERT
- ParsiNLU-RC
- ParsTwiNER
- PERLEX
- Pquadro
- Sepahr-Danesh
- byte-pair encoding
- WordPiece
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →