PulseAugur
EN
LIVE 05:39:18

New Mizo ASR system fine-tuned with Whisper and SraVaani models

Researchers have developed a new Automatic Speech Recognition (ASR) system for the Mizo language, a low-resource language, by collecting and curating over 17 hours of speech data. They fine-tuned three Whisper multilingual models and the SraVaani 1.0 Indic multilingual model. The Whisper-large-v3 model achieved the lowest word error rate (WER) of 18.08% conventionally, and 7.22% with morphology-aware evaluation. While SraVaani 1.0 showed a higher initial WER, fine-tuning significantly improved its performance, demonstrating the effectiveness of adapted models for unseen languages. AI

IMPACT Improves accessibility of speech technology for low-resource languages, potentially enabling new applications for Mizo speakers.

RANK_REASON The cluster describes a research paper detailing the creation and evaluation of a speech corpus and fine-tuning of ASR models for a low-resource language. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New Mizo ASR system fine-tuned with Whisper and SraVaani models

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Priyankoo Sarmah, Sanasam Ranbir Singh, Lalhmingmawia ·

    A Speech Corpus for Mizo Automatic Speech Recognition: Whisper and SraVaani 1.0 Fine-Tuning with Morphology-Aware Evaluation

    arXiv:2608.19361v1 Announce Type: new Abstract: This study reports the development of an Automatic Speech Recognition (ASR) system in Mizo, a low-resource language. The development included collecting 17.62 hours of speech data, curating it, and fine-tuning the Mizo ASR system wi…