Researchers have developed a novel Face-to-Speech (F2S) framework capable of generating plausible voices from static facial images, addressing the limitation of text-to-speech (TTS) systems that require reference audio. This system utilizes a lightweight Face Adapter to align facial recognition features with the style space of a frozen StyleTTS 2 model. Evaluations on the LRS3 dataset demonstrated highly natural synthesized speech, consistent voice generation, and a face-to-voice retrieval accuracy above chance. Notably, an English-trained adapter also produced fluent Spanish speech, suggesting the face-to-style mapping is largely language-agnostic. AI
IMPACT This research could enable new applications in historical reconstruction, virtual avatars, and accessibility tools by allowing voice generation from visual data.
RANK_REASON The cluster contains an academic paper detailing a new technical approach. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Face Adapter
- Hugging Face
- Jose A. Gonzalez-Lopez
- LRS3
- StyleTTS 2
- TED talk
- Zero-Shot Face-to-Speech Synthesis
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →