LibriSpeech
PulseAugur coverage of LibriSpeech — every cluster mentioning LibriSpeech across labs, papers, and developer communities, ranked by signal.
4 day(s) with sentiment data
-
New HearInContext benchmark tests speech recognition's grasp of implicit context
Researchers have introduced HearInContext, a new benchmark designed to evaluate the ability of speech recognition systems to understand implicit context. This benchmark, which includes 3,764 semantic test cases focusing…
-
Pruning Whisper-small improves ASR accuracy by acting as a regularizer
Researchers have demonstrated that neural network pruning can act as a regularization technique for Automatic Speech Recognition (ASR) systems, rather than just a method for compression. By analyzing the sensitivity of …
-
LLM merging enhances ASR performance without added computational cost
Researchers have developed a novel method for integrating large language models (LLMs) into automatic speech recognition (ASR) systems without increasing computational costs. This approach involves merging the LLMs dire…
-
New TTS models achieve lower latency and improved speech-text alignment · 4 sources tracked
Researchers have developed new methods for improving text-to-speech (TTS) systems, focusing on achieving lower latency and better alignment between text and speech. CTC-TTS utilizes a CTC-based aligner and a bi-word int…
-
New method significantly reduces hallucinations in Whisper ASR model
Researchers have developed a novel method to reduce "hallucinated transcripts" generated by the Whisper automatic speech recognition model. This training-free, inference-time technique projects decoder activations to su…
-
New SSRR loss boosts neural audio codec intelligibility and speed
A new research paper introduces a self-supervised representation reconstruction (SSRR) loss for neural audio codecs, aiming to improve intelligibility and reduce latency. This method accelerates training, allowing compe…
-
New benchmark evaluates AI speech model adaptation to accents
Researchers have developed a new benchmark called ABX- Accent to evaluate how well unsupervised speech models can adapt to different accents. The benchmark, based on the AESRC dataset, includes 10 English accents and a …
-
Hugging Face quantifies ASR benchmark optimization, finding models reproduce errors
Researchers at Hugging Face have developed new methods to quantify "benchmark optimization" or "benchmaxxing" in speech recognition models. Their study evaluated 11 open-source ASR models and found that several high-sco…
-
CPU Voice Cloning Benchmark: Pocket TTS, Kokoro, Audio8, XTTS v2 Compared
A practical benchmark evaluated four voice cloning models (Pocket TTS, Kokoro, Audio88 & Yassin, and XTTS v2) on CPU performance. The evaluation focused on speaker similarity, naturalness, intelligibility, and latency, …
-
Audio model research details compute-performance tradeoffs
A new research paper explores the optimal allocation of computational resources for audio models, focusing on Automatic Speech Recognition (ASR) and Speech Emotion Recognition (SER). The study introduces a framework tha…
-
CuteTTS system enhances speech synthesis with continuous autoregressive modeling
Researchers have developed CuteTTS, a novel text-to-speech system designed for efficient and high-quality voice synthesis. This system utilizes continuous autoregressive modeling with variational auto-encoder latents an…
-
New LILAC codec achieves idempotent audio tokenization
Researchers have developed LILAC, a novel neural speech codec designed to be idempotent, meaning it can decode and re-encode audio without altering the original token stream. This is a significant improvement over exist…
-
New Latent Softmax improves multilingual ASR data efficiency
Researchers have developed a new output layer called Latent Softmax for multilingual automatic speech recognition (ASR) systems. This method aims to improve data efficiency by better handling the differing supervision g…
-
KRAFTON releases bilingual speech model A.X K2 Raon-Speech
KRAFTON has released A.X K2 Raon-Speech, a bilingual English/Korean speech language model with 21.2 billion total parameters and 3.5 billion active parameters. This multimodal model integrates a text backbone from SK Te…
-
New diffusion model offers parallel speech transcription
Researchers have developed a novel approach to automatic speech recognition using a frozen discrete-diffusion language model, deviating from traditional autoregressive decoders. This new method refines entire transcript…
-
New COALA framework boosts speech recognition with contextual biasing
Researchers have developed COALA, a novel framework designed to improve automatic speech recognition (ASR) systems by integrating external knowledge. COALA enhances speech-augmented language models (SLMs) by mapping lat…
-
Apple and Cohere advance ASR with specialized and efficient models
Apple's Machine Learning Research team has developed a new approach to automatic speech recognition (ASR) error correction using compact seq2seq models. These models, trained on real and synthetic ASR errors, significan…
-
User seeks help implementing Calm TTS paper, facing voice cloning issues
A user is seeking assistance with implementing the Calm text-to-speech model described in a research paper. They have encountered difficulties in replicating the model's performance, experiencing issues with generating …
-
kandinskylab releases KVAE-Audio, a high-fidelity audio autoencoder
KVAE-Audio, a new continuous, full-band audio autoencoder, has been released by kandinskylab. This model effectively compresses raw audio waveforms into compact latents and reconstructs them with high fidelity across sp…
-
New neural diarization model excels on low-resource Nepali-Hindi speech
Researchers have developed a new approach to speaker diarization, the process of identifying who spoke when in an audio recording, specifically for low-resource languages like Nepali-Hindi. They trained two neural netwo…