PulseAugur
EN
LIVE 07:31:28

Open-source ASR models released for 27 African languages

Researchers have introduced DONDO, a collection of open-source automatic speech recognition (ASR) base models specifically designed for African languages. These models are built upon the w2v-BERT 2.0 architecture and include twenty-one monolingual and five multilingual models covering twenty-seven language varieties across several African nations. The models were fine-tuned using religious texts and a novel two-step fine-tuning procedure, achieving competitive word error rates and enabling a single multilingual model to perform well across multiple languages. AI

IMPACT This release provides valuable open-source tools for ASR research and development in underrepresented languages, potentially accelerating AI adoption in these regions.

RANK_REASON Release of open-source speech recognition models for multiple languages, detailed in an arXiv paper. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Open-source ASR models released for 27 African languages

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Paul Azunre ·

    DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages

    arXiv:2607.21540v1 Announce Type: new Abstract: We present DONDO, a family of open, permissively licensed automatic speech recognition (ASR) base models for African languages, built on the w2v-BERT 2.0 self-supervised speech encoder. DONDO comprises twenty-one monolingual models …