Hubert
PulseAugur coverage of Hubert — every cluster mentioning Hubert across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
Author criticizes "AI Nobel Prize" claims and AI inevitability
The author expresses frustration with the notion of an "AI Nobel Prize," arguing that it misattributes achievements to the AI itself rather than the human creators of the technology. They also dismiss the idea that AI i…
-
New SemDAC method boosts speech compression with semantic conditioning
Researchers have developed SemDAC, a novel neural speech compression method that prioritizes semantic content over waveform fidelity. By incorporating hierarchical semantic conditioning derived from HuBERT features, Sem…
-
Speech foundation models learn word representations beyond phonetics, study finds
A new research paper investigates whether self-supervised speech foundation models, such as HuBERT and wav2vec 2.0, truly learn word representations beyond just phonetic content. The study found that while these models …
-
Gemini 2.5-Flash leads multimodal rapport estimation in real-world HRI study
A new arXiv paper explores multimodal rapport estimation in real-world Human-Robot Interaction (HRI) settings, moving beyond controlled lab environments. Researchers found that zero-shot Large Language Models (LLMs) per…
-
New parameter-free method evaluates few-shot learning for elephant vocalizations
Researchers have developed a parameter-free method for evaluating few-shot learning in elephant vocalization classification. This approach uses nearest-centroid classification on fixed acoustic embeddings, comparing its…
-
AI model screens Parkinson's disease using face and voice without labels
Researchers have developed a novel method for screening Parkinson's disease using only facial expressions and voice analysis, eliminating the need for direct PD labels. This approach leverages frozen pretrained encoders…
-
New PINT method distills invariant linguistic content from speech
Researchers have developed a new method called PINT (Parallel Invariant Tokenization) to improve speech tokenization by focusing on the linguistic content that remains consistent across different utterances. This techni…
-
CF-Net uses multimodal fusion for ambivalence and hesitancy recognition
Researchers have developed CF-Net, a deep multimodal network designed to recognize ambivalence and hesitancy in videos. This network utilizes frozen SigLIP2, HuBERT, and DistilBERT backbones to process visual, audio, an…
-
ai-sage releases GigaAM Multilingual speech models
ai-sage has released GigaAM Multilingual, a family of Conformer-based foundation models. These models, available in 220M and 600M parameter variants, have been pre-trained on over 2 million hours of speech data spanning…
-
AI systems advance ambivalence and hesitancy recognition in video analysis · 8 sources tracked
Researchers have developed advanced methods for recognizing ambivalence and hesitancy in videos, participating in the 11th ABAW Challenge. One approach, the HSEmotion team's system, utilizes multi-task learning with fro…
-
TTS evaluation confounded by ASR family alignment, new ensembles proposed
Researchers have identified a significant confound in evaluating text-to-speech (TTS) systems using automatic speech recognition (ASR) verifiers. The apparent quality of these verifiers is heavily influenced by the ASR …
-
New syllabic tokenizer improves speech understanding by disentangling speaker identity
Researchers have developed a novel speaker-disentangled syllabic tokenizer that improves unsupervised syllabic tokenization by regressing speaker-perturbed representations toward clean targets within fixed-length chunks…
-
New evaluation set disentangles speaker and language effects in cross-lingual verification
Researchers have developed a new evaluation set for cross-lingual speaker verification (SV) systems, focusing on Iberian languages. This setup allows for the analysis of cross-lingual SV under consistent speaker identit…
-
BabyHuBERT model improves speaker segmentation in child speech recordings
Researchers have developed BabyHuBERT, a new self-supervised speech model specifically trained on multilingual, child-centered long-form recordings. This model aims to improve the segmentation of speakers in recordings …
-
MauBERT paper introduces multilingual phonetic representations for speech models
Researchers have developed MauBERT, a multilingual extension of the HuBERT self-supervised learning model. By incorporating articulatory features and a phonetic-to-articulatory mapping across 55 languages, MauBERT learn…
-
WavLM advances vocal effort classification with data augmentation
Researchers have advanced speaker-based vocal effort classification by utilizing the WavLM model, outperforming previous approaches like Wav2Vec2 and HuBERT. To combat data scarcity, they systematically studied various …
-
Speech models encode child age/gender in early layers, study finds
Researchers have analyzed how well self-supervised learning (SSL) models capture age and gender information in children's speech. The study focused on four models: Wav2Vec2, HuBERT, Data2Vec, and WavLM, examining their …
-
Transformer models show improved accuracy for Quranic ASR
Researchers have conducted a comparative study on pretrained Transformer models for Quranic Automatic Speech Recognition (ASR), aiming to reduce high Word Error Rates (WER) on user-recited verses. The study fine-tuned m…
-
New LM-SPT method enhances speech tokenization for better language model alignment
Researchers have developed LM-SPT, a novel method for speech tokenization that aims to improve the alignment between speech and language models. Unlike previous approaches that directly distill features or use pooling, …
-
New dataset enhances AI detection of deepfake audio with linguistic cues
Researchers have introduced Linguistically Augmented Audio Speech Data (LinguAS), a new dataset designed to combat the rise of deepfaked audio. LinguAS includes over 800 audio samples, both genuine and fake, annotated w…