AudioCaps
PulseAugur coverage of AudioCaps — every cluster mentioning AudioCaps across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New encoder-free audio captioning model CARD reduces inference costs
Researchers have developed CARD, a novel encoder-free audio captioning model that significantly reduces inference costs by removing the audio encoder. The model distills knowledge from a pretrained audio teacher, CLAP-H…
-
SwiftAudio uses caption-only distillation for efficient text-to-audio generation
Researchers have developed SwiftAudio, a novel one-step text-to-audio diffusion model that bypasses the need for paired audio data during distillation. This approach utilizes only text captions and a pre-trained diffusi…
-
FoleyGenEx framework unifies video-to-audio generation with advanced controls
Researchers have introduced FoleyGenEx, a novel framework for unified video-to-audio generation that addresses limitations in existing methods. FoleyGenEx integrates multi-modal control, frame-level temporal alignment, …
-
New method compresses audio tokens for language models
Researchers have developed a new method called Local Temporal Bipartite Merging (LTBM) to compress audio tokens in audio-language models. This training-free approach merges similar nearby audio tokens within a temporal …