Qwen3 ASR
PulseAugur coverage of Qwen3 ASR — every cluster mentioning Qwen3 ASR across labs, papers, and developer communities, ranked by signal.
5 day(s) with sentiment data
-
Study reveals ASR-roundtrip evaluation masks TTS reading errors
A new study published on arXiv highlights the limitations of ASR-roundtrip evaluation for assessing Chinese news text-to-speech (TTS) systems. The research found that this method can mask context- and convention-depende…
-
New audio-visual segmentation system wins LSVOS Challenge
Researchers have developed a novel system for audio-visual segmentation, which involves identifying and segmenting objects described in spoken language within a video. The system first transcribes speech to text using Q…
-
Open-source iOS app enables offline AI models on iPhone
An open-source iOS application called LiveTranscriber has been developed to run various speech and language models entirely on-device, enabling offline functionality on iPhones. The app supports models such as Whisper f…
-
Self-hosting open-source speech-to-text models incurs hidden costs
Self-hosting open-source speech-to-text models like Whisper Large V3, Qwen3 ASR, and NVIDIA's Parakeet and Canary can appear free initially, but the total cost of ownership is significant. Beyond the model weights, user…
-
AssemblyAI touts Universal-3.5 Pro over Qwen3-ASR for production speech-to-text
AssemblyAI has compared its Universal-3.5 Pro speech-to-text model against Alibaba's Qwen3-ASR, highlighting the advantages of its proprietary solution for production environments. While Qwen3-ASR is recognized as a cap…
-
AssemblyAI: Self-hosting AI models costs more than managed APIs
AssemblyAI argues that while self-hosting open-source speech models like Whisper or Qwen3-ASR on platforms such as Baseten, Modal, or Fireworks may seem cost-effective on paper, the total cost of ownership is often high…
-
r/LocalLLaMA users seek recommendations for ASR and TTS models
This Reddit post on r/LocalLLaMA asks for recommendations on Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) models. The original poster is currently using older models like Whisper and Kokoro with koboldcpp…
-
Voice assistant uses CPU for ASR/TTS, freeing GPU for LLM
A user has successfully implemented a voice assistant that runs speech recognition and text-to-speech models on a CPU, freeing up the GPU for the primary LLM. They tested the Qwen3-ASR and Kokoro-TTS ONNX models on a 20…
-
Hugging Face Transformers v5.13.0 adds KimiK, Xiaomi MiMo, NVIDIA, and Qwen ASR models
Hugging Face's Transformers library has released version 5.13.0, introducing several new open-source models. This update includes Kimi K2.5, a multimodal agentic model from KimiK, designed for advanced coding and autono…
-
New research advances ASR for dysarthric speech and synthetic data use · 4 sources tracked
Researchers are exploring new methods to improve automatic speech recognition (ASR) systems. One study details how fine-tuning the Whisper model with personalized data significantly reduced word error rates for dysarthr…
-
audio.cpp framework offers faster audio model inference
A new C++ inference framework called audio.cpp has been developed, built on top of ggml, to run various audio models including TTS, ASR, and voice conversion. The framework aims to consolidate multiple audio models into…
-
New compact TTS models Inflect-Micro-v2 and Inflect-Nano-v2 released
Owen has independently developed and funded Inflect-Micro-v2 and Inflect-Nano-v2, two new text-to-waveform speech synthesis models. These models are designed for local, fixed-voice English TTS, prioritizing either quali…
-
Google's Eloquent dictation app fails benchmarks due to frequent transcription errors
A user attempting to benchmark Google's new on-device dictation app, Eloquent, found it to be largely unusable due to frequent failures to transcribe audio. While the app's accuracy was competitive when it did function,…
-
Whisfusion uses masked diffusion for faster, more accurate speech recognition
Researchers have developed Whisfusion, a novel non-autoregressive system for automatic speech recognition (ASR) that utilizes masked diffusion models. This approach aims to match the accuracy of traditional autoregressi…
-
FormalASR converts spoken Chinese to formal text end-to-end
Researchers have developed FormalASR, a novel end-to-end system designed to convert spoken Chinese directly into formal written text. This approach bypasses the need for a separate post-editing step by an LLM, reducing …
-
Omi Health releases open-weight medical ASR model
Omi Health founder has released Omi Med STT v1, a fine-tuned version of NVIDIA's Parakeet TDT 0.6B model for medical Automatic Speech Recognition (ASR). This open-weight model is designed to run locally on devices, ensu…
-
New methods enhance simultaneous speech translation with decoder-only LLMs
Researchers are developing new methods for simultaneous speech translation, focusing on decoder-only large language models. One approach, AlignAtt4LLM, adapts attention mechanisms for these models to improve translation…
-
OpenBrief launches as local-first AI video summarizer
OpenBrief is a new open-source, local-first desktop application designed to help users process video and audio content. It allows users to import media, extract transcripts, generate summaries, and even chat with the co…
-
FormalASR system converts spoken Chinese to formal text end-to-end
Researchers have developed FormalASR, a novel end-to-end system designed to directly convert spoken Chinese into formal written text. This approach bypasses the need for a separate large language model (LLM) for post-ed…
-
Tech entrepreneur uses AI to manage home data migration and smart devices
A tech enthusiast and entrepreneur detailed his experience integrating AI into his home, starting with migrating his digital life to a new MacBook Pro. He utilized Claude Code, an AI assistant, to manage the complex tra…