Researchers have developed a synthetic Bengali speech dataset tailored for telecom customer care applications. This dataset comprises 10,000 audio-text pairs, totaling approximately 26.82 hours of speech, and is publicly available on Hugging Face under a CC-BY-4.0 license. The synthetic speech was generated using OmniVoice, with a focus on voice cloning and normalization for Automatic Speech Recognition (ASR) training. Evaluation using a fine-tuned Whisper ASR model indicates strong text-audio consistency, with low Word Error Rate (WER) and Character Error Rate (CER). AI
IMPACT This dataset could improve the performance of speech recognition systems in specific domains like telecom customer care.
RANK_REASON The item describes the creation and evaluation of a synthetic speech dataset, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Bangla
- bengaliAI/tugstugi_bengaliai-regional-asr_whisper-medium
- Creative Commons Attribution 4.0 International
- Hugging Face
- OmniVoice
- Whisper
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →