Whisper-medium
PulseAugur coverage of Whisper-medium — every cluster mentioning Whisper-medium across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New methods enhance Large Audio-Language Models via encoder selection and contrastive decoding · 2 sources tracked
Researchers have developed two novel methods to improve the performance of Large Audio-Language Models (LALMs). The first, CUES (Correlation-Guided Encoder Selection), uses a lightweight heuristic to select optimal enco…
-
Speech-to-SFT pipeline ablation reveals data quality gains don't always boost downstream performance
A new research paper published on arXiv details a factorial ablation study of a speech-to-SFT pipeline, investigating the impact of different refinement stages on data quality and downstream model performance. The study…
-
Whisper ASR Models Adapted for Multilingual Medical Use
Researchers have analyzed how multilingual medical adaptation affects the internal representations of Whisper models. The study compared various fine-tuning strategies across different Whisper model sizes, finding that …
-
New CASA system uses LLMs for interpretable speaking assessment
Researchers have developed CASA, a new system for automatic speaking assessment that uses a combination of the Whisper-medium and Qwen3.5-2B large language models. CASA achieves state-of-the-art performance with improve…
-
New benchmark compares multilingual models for Nepali ASR
A new study benchmarks six multilingual pre-trained models for Nepali Automatic Speech Recognition (ASR) using a standardized fine-tuning protocol. The research found that Whisper-Large-v3-Turbo and IndicWav2Vec perform…
-
Zero-shot voice cloning enhances dysarthric ASR model training
Researchers have explored zero-shot voice cloning as a method to augment datasets for automatic speech recognition (ASR) systems trained on dysarthric speech. By cloning speakers from the TORGO dataset using Higgs Audio…
-
ASR models evaluated on Dutch child speech, Whisper-medium leads
A new study published on arXiv evaluates the performance of nine state-of-the-art Automatic Speech Recognition (ASR) models, including Whisper, Parakeet, and Wav2Vec2, on Dutch child speech datasets. The fine-tuned Whis…