ECAPA-TDNN
PulseAugur coverage of ECAPA-TDNN — every cluster mentioning ECAPA-TDNN across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
Audio attribution models falter after audio transcoding, study finds
A new research paper published on arXiv investigates the robustness of audio provenance attribution systems when audio is transcoded. The study found that while these systems perform well on clean benchmarks, their accu…
-
New SISER model improves speech emotion recognition with adversarial training
Researchers have developed SISER, a novel approach to speech emotion recognition that addresses data scarcity and speaker variability. By integrating the pre-trained wav2vec 2.0 model for feature extraction and an ECAPA…
-
New method enables voice cloning in text-to-audio-video models
Researchers have developed a method to add voice cloning capabilities to text-to-audio-video (T2AV) generation models. By incorporating a single, zero-initialized linear layer and fine-tuning, T2AV models can be adapted…
-
New MERaLiON-GR model achieves state-of-the-art gender recognition across multiple languages
Researchers have developed MERaLiON-GR, a novel speech gender recognition model capable of classifying gender for both English and several Southeast Asian languages. This model is built upon MERaLiON-SpeechEncoder-2, a …
-
New framework improves speaker verification for non-verbal vocalizations
Researchers have developed a new framework for speaker verification that improves accuracy for non-verbal vocalizations (NVVs) while preserving performance on speech. The system combines frozen self-supervised features …
-
Speech-aware LLMs show weak speaker verification, new method improves performance
Researchers have developed a new method to evaluate and enhance the speaker verification capabilities of speech-aware Large Language Models (LLMs). Initial benchmarks revealed that current speech-aware LLMs exhibit weak…
-
Multimodal AI boosts classroom speaker identification accuracy
Researchers have developed a multimodal approach to speaker identification in K-12 classrooms, combining acoustic embeddings with Large Language Model (LLM) derived semantic context. This method significantly improved s…
-
AI Platform Combines Deepfake Detection with Blockchain Evidence
Researchers have developed a new platform called DeepFake Forensics AI, designed to combat the growing threat of synthetic media in legal and forensic settings. This system integrates multi-modal detection capabilities …
-
New method offers adaptive control over deep neural network sparsity
Researchers have developed an adaptive regularization method to better control sparsity in deep neural networks, addressing the challenge where traditional $\ell_1$ penalties indirectly influence sparsity rates. This ne…
-
Researchers develop new spoken language ID method using pre-trained models and margin loss
Researchers have developed a new method for spoken language identification using pre-trained models and margin-based losses. This approach enhances the ability of language representations to distinguish between language…
-
LASE model improves cross-script voice cloning by making embeddings language-uninformative
Researchers have developed LASE, a Language-Adversarial Speaker Encoder, to improve multilingual voice cloning. Standard encoders struggle to maintain speaker identity across different scripts, particularly when project…