WavLM-large
PulseAugur coverage of WavLM-large — every cluster mentioning WavLM-large across labs, papers, and developer communities, ranked by signal.
-
CPU Voice Cloning Benchmark: Pocket TTS, Kokoro, Audio8, XTTS v2 Compared
A practical benchmark evaluated four voice cloning models (Pocket TTS, Kokoro, Audio88 & Yassin, and XTTS v2) on CPU performance. The evaluation focused on speaker similarity, naturalness, intelligibility, and latency, …
-
PhonoQ enhances speech classification using audio-articulatory MRI data
Researchers have developed a method to improve the classification of speech based on audio and real-time MRI articulatory data. By incorporating representations from PhonoQ, an audio-based model trained on phonological …
-
Speech-LLM System Achieves High Accuracy in MLC-SLM Challenge
Researchers have developed a novel speech-LLM system for the 2nd MLC-SLM Challenge, focusing on automatic speech recognition and speaker diarization. Their system, which combines DiariZen-Large-s80 segmentation with CAM…
-
New framework improves recognition of human ambivalence and hesitancy
Researchers have developed a novel framework, SVF-CR, for recognizing subtle human emotional states like ambivalence and hesitancy by analyzing synchronized multimodal data. This approach refines visual and facial cues …
-
New research explores conversational timing and robust benchmarking for depression detection
Researchers are exploring new methods for detecting depression using conversational data. One study investigates the temporal dynamics of conversations, specifically the timing between clinician and participant turns, a…
-
Korean toddler pronunciation evaluated by AI
Researchers have developed an automated system to evaluate the pronunciation of Korean toddlers, addressing a gap in current assessment tools. The system utilizes neural speaker diarization and self-supervised learning …