Researchers have introduced MGhana-ST, a new speech translation dataset designed for four low-resource Ghanaian languages: Ga, Twi, Ewe, and Fante. The dataset, which includes paired audio and English translations with annotations for verbal and non-verbal events, is intended to advance research in African language speech technology. Experiments using the Whisper-small model revealed that flat multilingual training under data scarcity did not benefit all languages, with some showing declines in performance compared to monolingual training. AI
IMPACT This dataset and its analysis could inform future research and development of speech technology for underrepresented languages.
RANK_REASON The cluster describes a new academic paper introducing a dataset and experimental results on multilingual training trade-offs for low-resource languages. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →