Researchers have developed TutlAit v1, a new crowdsourced dataset for Moroccan Tamazight speech, addressing the scarcity of labeled audio data for this language. The dataset includes Arabic transcriptions and regional accent labels, collected via a custom web application. TutlAit v1 comprises over 13,000 audio files totaling approximately 21 hours, with a focus on Atlas and Souss varieties, and can be utilized for speech recognition, translation, and accent identification tasks. AI
IMPACT This dataset could significantly advance speech technology research for Moroccan Tamazight, enabling new applications in speech recognition and translation.
RANK_REASON The item describes a new academic dataset for a low-resource language, published on arXiv. [lever_c_demoted from research: ic=1 ai=1.0]
- Arabic
- Django 5
- django-rest-framework
- Elan
- Kabyle
- Modern Standard Arabic
- Mohamed-Amine Chadi
- Moroccan Tamazight
- PostgreSQL
- React 18
- SHA-256
- TutlAit
- TutlAit v1
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →