AudioSet
PulseAugur coverage of AudioSet — every cluster mentioning AudioSet across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New method speeds up Audio Spectrogram Transformer training with minimal accuracy loss
Researchers have developed a new method called SpecAugment-Patch Merging to improve the efficiency of training Audio Spectrogram Transformers (ASTs). This technique involves masking spectrograms at the patch level and t…
-
New research tackles catastrophic forgetting in AI sound classification
Researchers have investigated methods to prevent catastrophic forgetting in sound event classification tasks using incremental learning. The study analyzed architectural and regularization approaches, focusing on protec…
-
Audio-first triage slashes VLM calls for egocentric video captioning
Researchers have developed a novel audio-first approach for efficiently captioning long egocentric videos. This method prioritizes audio cues to decide which video segments are most relevant for analysis by a vision-lan…
-
New BAT system uses LLMs for spatial sound reasoning
Researchers have developed BAT, a system that combines a binaural acoustic scene analysis model with a large language model (LLM) to enable reasoning about spatial sounds. To facilitate this, they created a new dataset …
-
AV-JEPA model advances audio-visual self-supervised learning
Researchers have introduced AV-JEPA, a new self-supervised learning model that extends LeJEPA to handle both audio and visual data. This model utilizes an early-fusion Vision Transformer and modality dropout for masking…
-
Researchers Compare Token Representations Against CNNs for Bird Vocalization Detection
Researchers from DS@GT ARC explored token representations against supervised CNN backbones for the BirdCLEF+ 2026 challenge, which focuses on detecting animal vocalizations in soundscapes. They developed a baseline mode…
-
kandinskylab releases KVAE-Audio, a high-fidelity audio autoencoder
KVAE-Audio, a new continuous, full-band audio autoencoder, has been released by kandinskylab. This model effectively compresses raw audio waveforms into compact latents and reconstructs them with high fidelity across sp…
-
New probing method boosts audio SSL model evaluation
Researchers have developed a new method called binarized prototypical probes for evaluating audio self-supervised learning models. This technique addresses the information bottleneck caused by global pooling in existing…
-
New unsupervised method segments multilingual laughter in audio
Researchers have developed a new unsupervised method for segmenting acoustic laughter across multiple languages. This approach treats laughter detection as an anomaly detection problem on audio sequences, utilizing an I…