F5-TTS
PulseAugur coverage of F5-TTS — every cluster mentioning F5-TTS across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New attack reveals severe privacy risks in fine-tuned TTS models
Researchers have developed a new black-box membership inference attack (MIA) framework specifically designed for fine-tuned Text-to-Speech (TTS) models. This framework addresses challenges in query generation and repres…
-
New TTS pipeline enhances ASR systems with phoneme-based augmentation
Researchers have developed a unified pipeline for generating synthetic speech to improve automatic speech recognition (ASR) systems. This pipeline utilizes a multilingual text-to-speech (TTS) model, F5-TTS, with languag…
-
Author details five common pitfalls in fine-tuning F5-TTS models
The author details five critical mistakes made while fine-tuning the F5-TTS text-to-speech model for a new writing system. These errors, primarily related to loading the correct model weights and verifying vocabulary ch…
-
AI vocal synthesis for Russian faces quality, cost, and licensing hurdles
Generating realistic Russian vocals with AI presents significant challenges, as different models exhibit varying levels of performance with the same text. While some services struggle with pronunciation and stress, othe…
-
TTS evaluation confounded by ASR family alignment, new ensembles proposed
Researchers have identified a significant confound in evaluating text-to-speech (TTS) systems using automatic speech recognition (ASR) verifiers. The apparent quality of these verifiers is heavily influenced by the ASR …
-
New models unify speech and singing voice generation
Researchers have developed new unified models for generating human vocal audio, capable of producing both speech and singing. UniVoice uses a conditional flow matching approach, separating content, melody, and timbre to…