Researchers have developed a system named Aslema for the NADI 2026 Shared Task, focusing on intent recognition and slot filling for Tunisian Derja speech. The system demonstrated that fine-tuning large language models (LLMs) significantly outperforms zero-shot inference. By augmenting training data with LLM-generated synthetic utterances and voice cloning, Aslema achieved top rankings in slot filling and a strong performance in intent recognition on the test set. The team plans to release their scripts and synthetic dataset to foster further research in this domain. AI
IMPACT Demonstrates effective LLM-based data augmentation for low-resource languages, potentially improving speech recognition systems.
RANK_REASON The cluster contains an academic paper detailing a system for a specific NLP task and its performance, including a novel data augmentation technique. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →