Clotho
PulseAugur coverage of Clotho — every cluster mentioning Clotho across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New encoder-free audio captioning model CARD reduces inference costs
Researchers have developed CARD, a novel encoder-free audio captioning model that significantly reduces inference costs by removing the audio encoder. The model distills knowledge from a pretrained audio teacher, CLAP-H…
-
SwiftAudio uses caption-only distillation for efficient text-to-audio generation
Researchers have developed SwiftAudio, a novel one-step text-to-audio diffusion model that bypasses the need for paired audio data during distillation. This approach utilizes only text captions and a pre-trained diffusi…
-
New method compresses audio tokens for language models
Researchers have developed a new method called Local Temporal Bipartite Merging (LTBM) to compress audio tokens in audio-language models. This training-free approach merges similar nearby audio tokens within a temporal …