LibriSpeech
PulseAugur coverage of LibriSpeech — every cluster mentioning LibriSpeech across labs, papers, and developer communities, ranked by signal.
5 day(s) with sentiment data
-
CPU Voice Cloning Benchmark: Pocket TTS, Kokoro, Audio8, XTTS v2 Compared
A practical benchmark evaluated four voice cloning models (Pocket TTS, Kokoro, Audio88 & Yassin, and XTTS v2) on CPU performance. The evaluation focused on speaker similarity, naturalness, intelligibility, and latency, …
-
Audio model research details compute-performance tradeoffs
A new research paper explores the optimal allocation of computational resources for audio models, focusing on Automatic Speech Recognition (ASR) and Speech Emotion Recognition (SER). The study introduces a framework tha…
-
CuteTTS system enhances speech synthesis with continuous autoregressive modeling
Researchers have developed CuteTTS, a novel text-to-speech system designed for efficient and high-quality voice synthesis. This system utilizes continuous autoregressive modeling with variational auto-encoder latents an…
-
New LILAC codec achieves idempotent audio tokenization
Researchers have developed LILAC, a novel neural speech codec designed to be idempotent, meaning it can decode and re-encode audio without altering the original token stream. This is a significant improvement over exist…
-
New Latent Softmax improves multilingual ASR data efficiency
Researchers have developed a new output layer called Latent Softmax for multilingual automatic speech recognition (ASR) systems. This method aims to improve data efficiency by better handling the differing supervision g…
-
KRAFTON releases bilingual speech model A.X K2 Raon-Speech
KRAFTON has released A.X K2 Raon-Speech, a bilingual English/Korean speech language model with 21.2 billion total parameters and 3.5 billion active parameters. This multimodal model integrates a text backbone from SK Te…
-
New diffusion model offers parallel speech transcription
Researchers have developed a novel approach to automatic speech recognition using a frozen discrete-diffusion language model, deviating from traditional autoregressive decoders. This new method refines entire transcript…
-
New COALA framework boosts speech recognition with contextual biasing
Researchers have developed COALA, a novel framework designed to improve automatic speech recognition (ASR) systems by integrating external knowledge. COALA enhances speech-augmented language models (SLMs) by mapping lat…
-
Apple and Cohere advance ASR with specialized and efficient models
Apple's Machine Learning Research team has developed a new approach to automatic speech recognition (ASR) error correction using compact seq2seq models. These models, trained on real and synthetic ASR errors, significan…
-
User seeks help implementing Calm TTS paper, facing voice cloning issues
A user is seeking assistance with implementing the Calm text-to-speech model described in a research paper. They have encountered difficulties in replicating the model's performance, experiencing issues with generating …
-
kandinskylab releases KVAE-Audio, a high-fidelity audio autoencoder
KVAE-Audio, a new continuous, full-band audio autoencoder, has been released by kandinskylab. This model effectively compresses raw audio waveforms into compact latents and reconstructs them with high fidelity across sp…
-
New neural diarization model excels on low-resource Nepali-Hindi speech
Researchers have developed a new approach to speaker diarization, the process of identifying who spoke when in an audio recording, specifically for low-resource languages like Nepali-Hindi. They trained two neural netwo…
-
Hugging Face launches FFASR Leaderboard for real-world ASR benchmarking
Hugging Face and Treble Technologies have launched the FFASR Leaderboard, an open, community-driven benchmark for evaluating Automatic Speech Recognition (ASR) models in realistic far-field acoustic conditions. This new…
-
New ASR method InterAligner improves training stability and reduces errors
Researchers have developed a new method called InterAligner to improve the training stability and performance of Aligner-Encoder based Automatic Speech Recognition (ASR) models. This approach introduces an intermediate …
-
New research reveals CTC limitations in speech recognition, highlights linguistic model benefits
A new research paper explores the limitations of Connectionist Temporal Classification (CTC) in speech recognition systems. The study found that CTC's internal scoring methods struggle to improve accuracy beyond basic g…
-
New research tackles ASR challenges with synthetic speech, LLM optimization, and failure reduction
Researchers are developing advanced techniques to improve Automatic Speech Recognition (ASR) systems, particularly for challenging scenarios like code-switching and real-time applications. One paper proposes a code-mixi…
-
New NAR-MBR Decoding Boosts Speech Recognition Speed and Accuracy
Researchers have developed a new non-autoregressive decoding framework for speech recognition, termed NAR-MBR decoding. This method aims to improve the speed of speech recognition by generating output tokens in parallel…
-
Speech models compressed using parameter clustering
Researchers have developed a new method for compressing speech foundation models without requiring additional data or retraining. This approach utilizes channelwise clustering with k-means to achieve parameter compressi…
-
New model uses continuous space for speech recognition and translation
Researchers have introduced ELF-S2T, a novel approach to speech-to-text systems that operates in a continuous latent space rather than discrete text tokens. This model, built on the Embedded Language Flows (ELF) backbon…
-
New ASR methods tackle compute scaling and multilingual evaluation
Researchers are developing new methods to improve automatic speech recognition (ASR) systems. One approach, LARM, uses a depth-conditioned looped Transformer to allow for adjustable test-time computation, achieving perf…