PulseAugur
EN
LIVE 06:51:47
ENTITY word error rate

word error rate

PulseAugur coverage of word error rate — every cluster mentioning word error rate across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
10
29 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
6
22 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

8 day(s) with sentiment data

RECENT · PAGE 1/2 · 29 TOTAL
  1. TOOL · CL_239404 ·

    Whisper ASR tool aids Cantonese oral history transcription in New Zealand

    A new paper explores the application of Automatic Speech Recognition (ASR) tools, specifically Whisper, for multilingual oral history research. The study focused on Cantonese revitalization efforts in New Zealand, findi…

  2. RESEARCH · CL_233524 ·

    New metric OVMI standardizes speech BCI evaluation

    Researchers have introduced a new metric called open-vocabulary mutual information (OVMI) to standardize the evaluation of speech brain-computer interfaces (BCIs). This metric addresses the challenge of comparing BCIs t…

  3. TOOL · CL_229158 ·

    New batching method boosts speech transcription accuracy and speed

    A new paper introduces Context-Aware Interleaved Batching, a method designed to improve the accuracy and efficiency of speech transcription. This technique addresses limitations in existing systems like WhisperX by main…

  4. TOOL · CL_229042 ·

    New VoiceCodeBench benchmark evaluates ASR accuracy for structured tokens

    A new benchmark called VoiceCodeBench has been introduced to evaluate the accuracy of automatic speech recognition (ASR) systems in recovering specific structured tokens, such as identifiers and measured quantities. Tra…

  5. TOOL · CL_224404 ·

    Hugging Face adds Global South languages to Open ASR Leaderboard

    Hugging Face has expanded its Open ASR Leaderboard by introducing two new evaluation sets, Monsoon en-IN and Monsoon hi-IN, to better represent languages from the Global South. These additions aim to address the known i…

  6. TOOL · CL_212066 ·

    New Mizo ASR system fine-tuned with Whisper and SraVaani models

    Researchers have developed a new Automatic Speech Recognition (ASR) system for the Mizo language, a low-resource language, by collecting and curating over 17 hours of speech data. They fine-tuned three Whisper multiling…

  7. TOOL · CL_209287 ·

    AssemblyAI proposes Missed Entity Rate to improve speech-to-text accuracy

    AssemblyAI argues that traditional word error rate (WER) metrics for speech-to-text systems are insufficient because they fail to account for the importance of specific entities like names, numbers, and medical terms. T…

  8. RESEARCH · CL_210255 ·

    Whisper ASR Models Adapted for Multilingual Medical Use

    Researchers have analyzed how multilingual medical adaptation affects the internal representations of Whisper models. The study compared various fine-tuning strategies across different Whisper model sizes, finding that …

  9. COMMENTARY · CL_207289 ·

    Custom speech recognition models rarely outperform universal AI, says AssemblyAI

    AssemblyAI argues that custom speech recognition models are often unnecessary and can even be less accurate than modern universal models. While custom models were once the standard for improving accuracy in specialized …

  10. TOOL · CL_197189 ·

    AssemblyAI introduces cpWER to accurately measure speaker diarization accuracy

    AssemblyAI has introduced a new metric called cpWER (concatenated minimum-permutation word error rate) to more accurately measure the performance of speaker diarization in speech-to-text systems. Unlike traditional word…

  11. TOOL · CL_167715 ·

    Vision-Language Models Show Subtle Hallucinations in Historical Document OCR

    A new research paper analyzes the performance of vision-language models (VLMs) in transcribing historical documents, finding that while they outperform traditional optical character recognition (OCR) systems on standard…

  12. TOOL · CL_165096 ·

    Voice cloning enhances paralinguistic tasks and cross-lingual clinical speech analysis

    A new research paper explores the use of voice cloning for data augmentation in paralinguistic tasks, particularly for clinical applications where labeled data is scarce. The study benchmarks eight voice cloning models,…

  13. TOOL · CL_175926 ·

    Voice cloning models preserve paralinguistic signals for clinical speech tasks

    Researchers have evaluated eight voice cloning models to determine their effectiveness in preserving paralinguistic signals for speech tasks, particularly in clinical settings where labeled data is scarce. The study fou…

  14. TOOL · CL_155784 ·

    AssemblyAI details advanced speech recognition evaluation beyond WER

    AssemblyAI has published a guide detailing advanced methods for evaluating speech-to-text models, moving beyond the traditional Word Error Rate (WER). The article highlights the limitations of WER, such as its inability…

  15. TOOL · CL_154450 ·

    Whisper model fine-tuned for robust Assamese speech recognition

    Researchers have developed a fine-tuned version of the Whisper model to improve Automatic Speech Recognition (ASR) for the Assamese language. The fine-tuned model, trained on the Mozilla Common Voice 24.0-Assamese corpu…

  16. TOOL · CL_144437 ·

    AssemblyAI details best practices for production voice agents

    AssemblyAI has published a series of blog posts detailing best practices for building production-ready voice agents. The articles emphasize the importance of robust telemetry and diagnostic pipelines to catch regression…

  17. TOOL · CL_149538 ·

    New RLHF framework improves Vietnamese translation of historical manuscripts

    Researchers have developed a new multimodal Reinforcement Learning from Human Feedback (RLHF) framework to translate historical Han-Nom manuscripts into modern Vietnamese. This approach leverages both the visual informa…

  18. RESEARCH · CL_135148 ·

    New GRPO method boosts synthetic speech ASR performance

    Researchers have developed a new method called Group Relative Policy Optimization (GRPO) to improve automatic speech recognition (ASR) models, particularly when trained on synthetic speech. This reinforcement learning a…

  19. RESEARCH · CL_135192 ·

    Qwen-ASR-1.7B adapted for multilingual two-speaker speech recognition · 2 sources tracked

    Researchers have developed a system for the MLC-SLM 2026 Challenge that adapts the Qwen3-ASR-1.7B model for multilingual, two-speaker conversational speech. The system integrates a speaker diarization front end with the…

  20. RESEARCH · CL_115291 ·

    New pipeline enhances ASR robustness, cutting word error rate by 55%

    Researchers have developed a novel dual-gate diagnostic pipeline to enhance the robustness of Automatic Speech Recognition (ASR) systems against adversarial and benign perturbations. This pipeline, featuring a Two-Sided…