PulseAugur
EN
LIVE 06:35:06

New Framework Generates Speech from Facial Images

Researchers have developed a novel Face-to-Speech (F2S) framework capable of generating plausible voices from static facial images, addressing the limitation of text-to-speech (TTS) systems that require reference audio. This system utilizes a lightweight Face Adapter to align facial recognition features with the style space of a frozen StyleTTS 2 model. Evaluations on the LRS3 dataset demonstrated highly natural synthesized speech, consistent voice generation, and a face-to-voice retrieval accuracy above chance. Notably, an English-trained adapter also produced fluent Spanish speech, suggesting the face-to-style mapping is largely language-agnostic. AI

IMPACT This research could enable new applications in historical reconstruction, virtual avatars, and accessibility tools by allowing voice generation from visual data.

RANK_REASON The cluster contains an academic paper detailing a new technical approach. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New Framework Generates Speech from Facial Images

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Carlos Mu\~noz-Romero, Jose A. Gonzalez-Lopez ·

    Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model

    arXiv:2607.26742v1 Announce Type: cross Abstract: Zero-shot text-to-speech (TTS) clones a voice from a short audio prompt, but this reliance on reference audio is a barrier when only visual information is available, e.g. for historical figures or video-game characters. In this wo…