PulseAugur
EN
LIVE 23:04:33

Scenema Audio integrates with ComfyUI, enabling expressive TTS on 8GB VRAM

Scenema Audio, a text-to-speech model capable of zero-shot voice cloning and performing expressive speech with inline stage directions, has been released as a custom node for ComfyUI. This new integration allows the model to run on systems with as little as 8GB of VRAM, making it more accessible for self-hosting. The release simplifies prompt formatting from XML to inline bracket cues and includes twelve preset voices, though users may need to accept Gemma 3 12B model licenses and provide an HF token for initial setup. AI

IMPACT Enhances creative workflows by making advanced generative audio tools more accessible on consumer hardware.

RANK_REASON New integration of an existing AI model into a popular creative workflow tool.

Read on r/StableDiffusion →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Scenema Audio integrates with ComfyUI, enabling expressive TTS on 8GB VRAM

COVERAGE [1]

  1. r/StableDiffusion TIER_2 English(EN) · /u/a__side_of_fries ·

    Scenema Audio Comes to ComfyUI, Runs on 8GB VRAM

    <table> <tr><td> <a href="https://www.reddit.com/r/StableDiffusion/comments/1vgfhrp/scenema_audio_comes_to_comfyui_runs_on_8gb_vram/"> <img alt="Scenema Audio Comes to ComfyUI, Runs on 8GB VRAM" src="https://external-preview.redd.it/dDY5d3RkeGhtbGhoMfDbbs74itHS_b4Z-x5HLPmm6FqiLRj…