Hubert
PulseAugur coverage of Hubert — every cluster mentioning Hubert across labs, papers, and developer communities, ranked by signal.
4 day(s) with sentiment data
-
AI model screens Parkinson's disease using face and voice without labels
Researchers have developed a novel method for screening Parkinson's disease using only facial expressions and voice analysis, eliminating the need for direct PD labels. This approach leverages frozen pretrained encoders…
-
New PINT method distills invariant linguistic content from speech
Researchers have developed a new method called PINT (Parallel Invariant Tokenization) to improve speech tokenization by focusing on the linguistic content that remains consistent across different utterances. This techni…
-
CF-Net uses multimodal fusion for ambivalence and hesitancy recognition
Researchers have developed CF-Net, a deep multimodal network designed to recognize ambivalence and hesitancy in videos. This network utilizes frozen SigLIP2, HuBERT, and DistilBERT backbones to process visual, audio, an…
-
ai-sage releases GigaAM Multilingual speech models
ai-sage has released GigaAM Multilingual, a family of Conformer-based foundation models. These models, available in 220M and 600M parameter variants, have been pre-trained on over 2 million hours of speech data spanning…
-
AI systems advance ambivalence and hesitancy recognition in video analysis · 8 sources tracked
Researchers have developed advanced methods for recognizing ambivalence and hesitancy in videos, participating in the 11th ABAW Challenge. One approach, the HSEmotion team's system, utilizes multi-task learning with fro…
-
TTS evaluation confounded by ASR family alignment, new ensembles proposed
Researchers have identified a significant confound in evaluating text-to-speech (TTS) systems using automatic speech recognition (ASR) verifiers. The apparent quality of these verifiers is heavily influenced by the ASR …
-
New syllabic tokenizer improves speech understanding by disentangling speaker identity
Researchers have developed a novel speaker-disentangled syllabic tokenizer that improves unsupervised syllabic tokenization by regressing speaker-perturbed representations toward clean targets within fixed-length chunks…
-
New evaluation set disentangles speaker and language effects in cross-lingual verification
Researchers have developed a new evaluation set for cross-lingual speaker verification (SV) systems, focusing on Iberian languages. This setup allows for the analysis of cross-lingual SV under consistent speaker identit…
-
BabyHuBERT model improves speaker segmentation in child speech recordings
Researchers have developed BabyHuBERT, a new self-supervised speech model specifically trained on multilingual, child-centered long-form recordings. This model aims to improve the segmentation of speakers in recordings …
-
MauBERT paper introduces multilingual phonetic representations for speech models
Researchers have developed MauBERT, a multilingual extension of the HuBERT self-supervised learning model. By incorporating articulatory features and a phonetic-to-articulatory mapping across 55 languages, MauBERT learn…
-
WavLM advances vocal effort classification with data augmentation
Researchers have advanced speaker-based vocal effort classification by utilizing the WavLM model, outperforming previous approaches like Wav2Vec2 and HuBERT. To combat data scarcity, they systematically studied various …
-
Speech models encode child age/gender in early layers, study finds
Researchers have analyzed how well self-supervised learning (SSL) models capture age and gender information in children's speech. The study focused on four models: Wav2Vec2, HuBERT, Data2Vec, and WavLM, examining their …
-
Transformer models show improved accuracy for Quranic ASR
Researchers have conducted a comparative study on pretrained Transformer models for Quranic Automatic Speech Recognition (ASR), aiming to reduce high Word Error Rates (WER) on user-recited verses. The study fine-tuned m…
-
New LM-SPT method enhances speech tokenization for better language model alignment
Researchers have developed LM-SPT, a novel method for speech tokenization that aims to improve the alignment between speech and language models. Unlike previous approaches that directly distill features or use pooling, …
-
New dataset enhances AI detection of deepfake audio with linguistic cues
Researchers have introduced Linguistically Augmented Audio Speech Data (LinguAS), a new dataset designed to combat the rise of deepfaked audio. LinguAS includes over 800 audio samples, both genuine and fake, annotated w…
-
Speech models generalize to recognize rare click consonants
Researchers investigated whether self-supervised speech models can accurately recognize uncommon speech sounds, specifically click consonants found in Khoisan languages. By fine-tuning models like Wav2Vec2 and HuBERT on…
-
AI model detects Parkinson's disease using multi-modal speech analysis
Researchers have developed a novel multi-branch deep learning framework designed to improve the detection of Parkinson's disease through speech analysis. This approach utilizes three distinct speech representations: Log…
-
GeMCL algorithm scales few-shot spoken word classification
Researchers have developed a new method called Generative Meta-Continual Learning (GeMCL) to improve few-shot spoken word classification. This approach allows a model to sequentially learn to distinguish between 1000 cl…
-
New NLP Models Tackle Dementia Detection in Filipino Speech
Researchers have developed a new approach to dementia detection using natural language processing, focusing on low-resource languages like Filipino. They created a bilingual dataset and evaluated several transformer mod…
-
Generative meta-learning shows minimal language impact on spoken word classification
Researchers have explored the effectiveness of generative meta-continual learning for spoken word classification across multiple languages. Their findings indicate that while multilingual models perform best, the perfor…