Yoruba
PulseAugur coverage of Yoruba — every cluster mentioning Yoruba across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
Study finds LLM data audits don't guarantee downstream utility for African NLP
A new study published on arXiv investigates the effectiveness of synthetic data selection methods for low-resource African languages. The research found that common proxies, which assume that data rated highly by an LLM…
-
New research tackles multilingual AI efficiency and capabilities
Researchers are developing new methods to improve the efficiency and capabilities of multilingual AI models. One study explores token merging for multilingual speech recognition, showing it can significantly reduce comp…
-
New evaluation method reveals ASR/audio LM struggles with English-Yoruba code-switching
A new research paper introduces a switch-aware evaluation method for automatic speech recognition (ASR) and audio language models (audio LMs) when processing code-switched speech, specifically focusing on English and Yo…
-
First City Monument Bank to offer AI advisory in local languages for farmers
First City Monument Bank is integrating AI to offer advisory services in Hausa, Yoruba, and Igbo languages. This initiative aims to support smallholder farmers. The bank's efforts are detailed in its Agritech Alumni Imp…
-
New method recovers AI safety for African languages without retraining
Researchers have developed a novel training-free method called Latent Space Refusal Anchoring (LSR-Anchoring) to improve safety in instruction-tuned AI models for low-resource African languages. This technique aims to r…
-
New AfriNLLB models offer efficient translation for 15 African languages
Researchers have developed AfriNLLB, a suite of lightweight translation models designed for African languages. These models are derived from the NLLB-200 600M architecture, which has been compressed through layer prunin…
-
Tokenization premiums create AI cost barriers for non-English languages · arXiv cs.CL
A new study published on arXiv introduces the Tokenization Equity Audit (TEA), a benchmark designed to measure disparities in how large language models tokenize different languages. The research found that semantically …
-
AI's impact on linguistic diversity in World Englishes debated in new papers · 3 sources tracked
Three recent academic papers explore the complex relationship between generative AI and linguistic diversity, particularly concerning World Englishes. The first paper discusses how AI tools can both democratize academic…
-
New speech synthesizer developed for Yoruba language
Researchers have developed TTSYoruba, a rule-based speech synthesizer for the Yoruba language. This system utilizes a hand-crafted phonological rule system and a recorded inventory of diphone units to convert tone-marke…
-
Developer builds AI with memory system, using specialized Qwen models
A developer has created a novel AI memory system named Àtúnbí, meaning "reborn" in Yoruba, designed to overcome the limitations of stateless AI conversations. This system prioritizes information based on importance and …
-
OmniVoice fine-tuned for Yoruba zero-shot voice cloning
A developer fine-tuned the OmniVoice text-to-speech model for the Yoruba language, a tonal language where precise pronunciation is critical for meaning. The process involved constructing a dataset by merging high-qualit…
-
NSFW Anime Illustrations Promoted on Mastodon
This cluster contains two Mastodon posts featuring anime-style illustrations of characters, Rikka Takarada and Yor. The posts include hashtags related to NSFW content, hentai, and AI-generated imagery. They also promote…
-
Study finds contrastive prompts boost African language NLI performance
A new study published on arXiv explores prompting strategies for Natural Language Inference (NLI) in low-resource African languages, specifically Swahili, Yoruba, and Hausa. Researchers evaluated five different promptin…
-
New Corpus Aims to Boost African Languages in Science
A new parallel corpus called AfriScience-MT has been developed to address the lack of scientific terminology in six African languages: Amharic, Hausa, Luganda, Northern Sotho, Yorùbá, and isiZulu. This corpus, created b…
-
New SBPN model boosts Nigerian language ASR via knowledge distillation
Researchers have developed a new multilingual Automatic Speech Recognition (ASR) framework called Sometin Beta Pass Notin (SBPN) to improve performance for Nigerian languages. The framework uses a two-stage knowledge di…
-
CURE-Med framework enhances multilingual medical reasoning in LLMs
Researchers have developed CURE-Med, a novel framework using curriculum-informed reinforcement learning to enhance multilingual medical reasoning in large language models. This approach integrates code-switching-aware s…