ESC-50
PulseAugur coverage of ESC-50 — every cluster mentioning ESC-50 across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New MADS descriptor set enhances audio analysis beyond spectral summaries
Researchers have developed MADS (Multi-view Acoustic Descriptor Set), a new 19-dimensional descriptor set designed to capture richer physical dynamics in audio signals beyond traditional spectral summaries. MADS encodes…
-
Curiosity-driven MoE framework enhances AI model stability and accuracy on edge devices
Researchers have developed a novel curiosity-driven quantized Mixture-of-Experts (MoE) framework designed for resource-constrained devices. This approach addresses challenges in maintaining accuracy under aggressive qua…
-
New research identifies transferable sentiment axis in LLMs across modalities
Researchers have identified a single internal direction within large language models that effectively tracks the sentiment of text, termed the valence axis (V-axis). This V-axis can be discovered using a minimal set of …
-
New framework evaluates sound effects generation systems
Researchers have developed a new framework for evaluating sound effects (SFX) generation systems, addressing the need for realistic audio that also maintains perceptual identity and allows for controllable variation. Th…
-
MJEPA architecture simplifies audio-visual learning with unified encoder
Researchers have introduced MJEPA, a novel architecture for audio-visual learning that utilizes a single, unified encoder for both modalities. This approach simplifies existing methods by employing a single predictive o…
-
MJEPA: Unified Audio-Visual Learning Architecture Unveiled
Researchers have introduced MJEPA, a novel joint-embedding predictive architecture designed for audio-visual learning. This approach utilizes a single, unified encoder for both modalities, simplifying the learning proce…
-
AudioPG uses synthetic data for efficient audio model pre-training
Researchers have developed AudioPG, a novel framework for pre-training audio models using procedurally generated synthetic data instead of real-world recordings. This approach significantly reduces training costs, curat…
-
New criterion predicts effectiveness of time-lagged spectral embeddings
Researchers have developed a new criterion to determine the applicability of training-free time-lagged spectral embeddings for multivariate time series. This criterion, based on stationarity and temporal coupling, predi…
-
Diffusion model advances zero-shot environmental sound classification
Researchers have developed a novel diffusion model for zero-shot environmental sound classification, a task that has historically struggled with poor performance. This new model generates synthetic embeddings for unseen…