Whisper Large V3
PulseAugur coverage of Whisper Large V3 — every cluster mentioning Whisper Large V3 across labs, papers, and developer communities, ranked by signal.
5 day(s) with sentiment data
-
New model DiaWhisper-DPO improves clinical interview transcription and role attribution
Researchers have developed DiaWhisper-DPO, an end-to-end model for transcribing clinical interviews and attributing utterances to either the clinician or patient. This model fine-tunes Whisper-large-v3 using LoRA and an…
-
Whisper adaptation improves Greek song lyric transcription
Researchers have developed a new method for Automatic Lyric Transcription (ALT) specifically for Greek songs, addressing the challenges posed by melodic and rhythmic variations. By adapting OpenAI's Whisper model, they …
-
BuzzASR releases 100+ language-specific speech models, outperforming Whisper
Researchers have developed BuzzASR, a suite of over 100 specialized speech recognition models fine-tuned for individual languages. These models are based on the Whisper architecture and significantly outperform the gene…
-
PEFT variants compared for dysarthric ASR in single-speaker study
Researchers have conducted a case study comparing seven Parameter-Efficient Fine-Tuning (PEFT) variants for per-patient Automatic Speech Recognition (ASR) systems designed for individuals with dysarthria. The study focu…
-
Groq API's model list includes non-chat models and hidden token limits
A review of the Groq API's model listing revealed that five of the fourteen advertised models are not capable of chat completions. These non-chat models include speech-to-text and text-to-speech variants, as well as a r…
-
Together AI ranks top open-source models by use case
Together AI has released a comparative analysis of top open-source AI models, categorizing them by their performance across various use cases. The analysis highlights models like Kimi K3, DeepSeek V4 Pro, and Qwen3.8 2.…
-
On-device AI system developed for breast cancer multidisciplinary team meetings
Researchers have developed an on-device AI system designed to assist in breast cancer multidisciplinary team meetings. This system utilizes open-source Automatic Speech Recognition (ASR) and Large Language Models (LLMs)…
-
New Mizo ASR system fine-tuned with Whisper and SraVaani models
Researchers have developed a new Automatic Speech Recognition (ASR) system for the Mizo language, a low-resource language, by collecting and curating over 17 hours of speech data. They fine-tuned three Whisper multiling…
-
Whisper ASR Models Adapted for Multilingual Medical Use
Researchers have analyzed how multilingual medical adaptation affects the internal representations of Whisper models. The study compared various fine-tuning strategies across different Whisper model sizes, finding that …
-
Free Whisper bots outperform paid services on noisy Russian audio
A recent independent review indicates that free Whisper-based bots can outperform paid services like Otter.ai for transcribing noisy Russian audio. While services like Adobe Enhance Speech and Cleanvoice AI offer audio …
-
Self-hosting open-source speech-to-text models incurs hidden costs
Self-hosting open-source speech-to-text models like Whisper Large V3, Qwen3 ASR, and NVIDIA's Parakeet and Canary can appear free initially, but the total cost of ownership is significant. Beyond the model weights, user…
-
AssemblyAI touts Universal-3.5 Pro over Whisper for production speech-to-text
AssemblyAI has published a comparison highlighting the advantages of its Universal-3.5 Pro model over OpenAI's Whisper Large-v3 for production speech-to-text applications. While Whisper is effective for clean audio and …
-
New AI model predicts Arabic speaker origin using continuous dialect space
Researchers have developed a novel regression-based method to predict the geographic origin of Arabic speakers by modeling dialectal variations as a continuous space. The approach utilizes a hierarchical neural network …
-
Audio separation harms zero-shot ASR performance, study finds
A new research paper investigates the counterintuitive finding that audio separation can degrade the performance of zero-shot Automatic Speech Recognition (ASR) systems. The study evaluated SAM-Audio as a preprocessing …
-
New Whisper-based system improves prosodic boundary detection in Brazilian Portuguese
Researchers have developed SAMPA, a new system for automatically segmenting prosodic boundaries in Brazilian Portuguese speech. This system is based on fine-tuning the Whisper large-v3 model, a significant advancement o…
-
NVIDIA unveils Audex, a unified audio-text LLM that preserves text intelligence
NVIDIA has introduced Audex (Nemotron-Labs-Audex-30B-A3B), a unified audio-text large language model capable of understanding and generating both audio and speech. Unlike many multimodal models that experience a decline…
-
Apple and Cohere advance ASR with specialized and efficient models
Apple's Machine Learning Research team has developed a new approach to automatic speech recognition (ASR) error correction using compact seq2seq models. These models, trained on real and synthetic ASR errors, significan…
-
ASR systems and humans struggle with Dutch dysarthric speech recognition
A new study published on arXiv compares the performance of human listeners and three advanced automatic speech recognition (ASR) systems—Whisper-large-V3, Google Chirp 3, and Omnilingual—in recognizing Dutch dysarthric …
-
New compact TTS models Inflect-Micro-v2 and Inflect-Nano-v2 released
Owen has independently developed and funded Inflect-Micro-v2 and Inflect-Nano-v2, two new text-to-waveform speech synthesis models. These models are designed for local, fixed-voice English TTS, prioritizing either quali…
-
LibriConvo corpus advances ASR and speaker diarization
Researchers have developed LibriConvo, a new synthetic conversational speech corpus designed to improve automatic speech recognition (ASR) and speaker diarization systems. The corpus was created by adapting the Speaker-…