PulseAugur
EN
LIVE 10:02:54
ENTITY Whisper

Whisper

PulseAugur coverage of Whisper — every cluster mentioning Whisper across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
31
142 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
18
60 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
TIMELINE
  1. 2026-08-27 research_milestone A study fine-tuned the Whisper model for automatic speech recognition in the Baniwa language, achieving notable error rate reductions. source
  2. 2026-06-09 research_milestone A study on fine-tuning OpenAI's Whisper for Swiss German ASR revealed improved performance and identified benchmark contamination issues. source
  3. 2026-05-12 research_milestone A new semi-supervised framework for speech confidence detection was proposed, achieving a Macro-F1 score of 0.751. source
SENTIMENT · 30D

13 day(s) with sentiment data

RECENT · PAGE 1/10 · 200 TOTAL
  1. RESEARCH · CL_259286 ·

    T-SANDHI improves Taiwanese Hokkien ASR by decoupling tone variations

    Researchers have developed T-SANDHI, a novel approach to improve automatic speech recognition (ASR) for low-resource languages like Taiwanese Hokkien. Unlike previous assumptions that tone sandhi is the primary challeng…

  2. TOOL · CL_256996 ·

    Whisper model adapted for police body camera audio transcription

    Researchers have developed a method to adapt OpenAI's Whisper model for transcribing audio from law enforcement body cameras. This adaptation uses Low-Rank Adaptation (LoRA) to improve accuracy in challenging acoustic e…

  3. TOOL · CL_256713 ·

    Open-source tools enable private, self-hosted voice AI assistants

    An engineer named Ravi Roy highlights the growing trend of building private, open-source voice AI assistants, moving away from proprietary cloud services. This approach offers enhanced data privacy, reduced costs, and g…

  4. TOOL · CL_254559 ·

    New method boosts low-resource language ASR using adapter stacking

    Researchers have developed a new method called Sequential Adapter Stacking to improve automatic speech recognition (ASR) for low-resource languages. This technique involves layering trainable target-language adapters on…

  5. TOOL · CL_254545 ·

    New streaming Thai speech recognition system offers low-latency and steerable vocabulary

    Researchers have developed a novel streaming Thai speech recognition system called Typhoon ASR Streaming, designed for low-latency applications. This system addresses the limitations of existing Whisper-based models, wh…

  6. TOOL · CL_254537 ·

    Neyshekar corpus released for Persian speech recognition

    A new open Persian read-speech corpus named Neyshekar has been released, containing over 62,000 recordings totaling nearly 100 hours. This corpus is designed to cover formal and informal Persian, named entities, and lon…

  7. RESEARCH · CL_252080 ·

    New research tackles multilingual AI efficiency and capabilities

    Researchers are developing new methods to improve the efficiency and capabilities of multilingual AI models. One study explores token merging for multilingual speech recognition, showing it can significantly reduce comp…

  8. TOOL · CL_251974 ·

    New method quantifies consonant contributions to word intelligibility

    Researchers have developed a novel method to quantify the contribution of consonants to word intelligibility using acoustic masking. This technique involves masking individual consonants within a word and assessing whet…

  9. COMMENTARY · CL_248156 ·

    Open-source multimodal models challenge GPT-4o on cost and performance · 2 sources tracked

    Open-source multimodal models are rapidly catching up to GPT-4o in performance and cost-effectiveness, with several models like Alibaba's Qwen2.5-VL and Mistral's Pixtral 12B offering competitive capabilities for tasks …

  10. TOOL · CL_247720 ·

    New methods improve multilingual video transcription accuracy

    Researchers have developed methods to improve speech transcription accuracy from videos across multiple languages, aiming to aid the creation of automated tools for cross-cultural understanding. By leveraging publicly a…

  11. TOOL · CL_247686 ·

    Whisper adaptation improves Greek song lyric transcription

    Researchers have developed a new method for Automatic Lyric Transcription (ALT) specifically for Greek songs, addressing the challenges posed by melodic and rhythmic variations. By adapting OpenAI's Whisper model, they …

  12. TOOL · CL_245603 ·

    New SemDAC method boosts speech compression with semantic conditioning

    Researchers have developed SemDAC, a novel neural speech compression method that prioritizes semantic content over waveform fidelity. By incorporating hierarchical semantic conditioning derived from HuBERT features, Sem…

  13. TOOL · CL_245290 ·

    AI cognitive screening models show significant bias against multilingual speakers

    A new study published on arXiv has identified a significant false-positive bias in AI models used for speech-based cognitive screening, particularly affecting multilingual individuals in the UK. The research found that …

  14. RESEARCH · CL_245281 ·

    NOPE-HYPE workflow enhances speech-to-text model robustness via simulation

    Researchers have developed NOPE-HYPE, a novel structured workflow designed to improve the robustness of speech-to-text models across varied acoustic environments. This workflow integrates a controllable simulator for ac…

  15. RESEARCH · CL_245231 ·

    BuzzASR releases 100+ language-specific speech models, outperforming Whisper

    Researchers have developed BuzzASR, a suite of over 100 specialized speech recognition models fine-tuned for individual languages. These models are based on the Whisper architecture and significantly outperform the gene…

  16. SIGNIFICANT · CL_241804 ·

    Xiaomi MiMo-V2.5: Open-weight omnimodal AI with 1M+ token context

    Xiaomi has released MiMo-V2.5, an open-weight omnimodal AI model designed to process audio, image, and video natively within a single context. While the model's text reasoning, coding, and JSON extraction capabilities h…

  17. TOOL · CL_240943 ·

    Open-source tools offer alternative to costly AI subscriptions

    A developer has outlined a suite of five open-source tools that can replace a costly monthly AI subscription, emphasizing control and cost savings. The proposed stack includes Ollama for local LLM inference, LiteLLM as …

  18. TOOL · CL_239692 ·

    Olud Pulse tracks open-source AI adoption: Open Sora, Whisper, pgvector lead categories · 3 sources tracked

    Olud Pulse, a tool that tracks the adoption of open-source AI, has released its latest scores. Open Sora leads in AI video generation, Whisper is highest in speech recognition and text-to-speech, and pgvector tops the v…

  19. TOOL · CL_239404 ·

    Whisper ASR tool aids Cantonese oral history transcription in New Zealand

    A new paper explores the application of Automatic Speech Recognition (ASR) tools, specifically Whisper, for multilingual oral history research. The study focused on Cantonese revitalization efforts in New Zealand, findi…

  20. TOOL · CL_239210 ·

    New method significantly reduces hallucinations in Whisper ASR model

    Researchers have developed a novel method to reduce "hallucinated transcripts" generated by the Whisper automatic speech recognition model. This training-free, inference-time technique projects decoder activations to su…