PulseAugur
EN
LIVE 14:04:18

FireRedTeam unveils unified audio model FireRedAudio and TTS3

FireRedTeam has released FireRedAudio, a 9-billion parameter audio language model capable of a wide range of tasks including automatic speech recognition, audio understanding, and various forms of text-to-speech (TTS). The model utilizes decoupled continuous representations for understanding and generation, allowing it to process and generate speech for recordings up to an hour long. Alongside FireRedAudio, they also introduced FireRedTTS3, a unified system for speech generation and editing that supports zero-shot voice cloning across 24 languages and 21 Chinese dialects, with an instruct variant for natural language control. AI

IMPACT This release offers a unified approach to audio processing and generation, potentially simplifying workflows for developers working with speech and audio data.

RANK_REASON Release of a new open-source audio language model and TTS system. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

FireRedTeam unveils unified audio model FireRedAudio and TTS3

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/pmttyji ·

    FireRedAudio & FireRedTTS3 by FireRedTeam - Huggingface

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vukj3m/fireredaudio_fireredtts3_by_fireredteam/"> <img alt="FireRedAudio &amp; FireRedTTS3 by FireRedTeam - Huggingface" src="https://preview.redd.it/sxn3p1m82rkh1.png?width=640&amp;crop=smart&amp;auto=webp&a…