Scenema Audio, a text-to-speech model capable of zero-shot voice cloning and performing expressive speech with inline stage directions, has been released as a custom node for ComfyUI. This new integration allows the model to run on systems with as little as 8GB of VRAM, making it more accessible for self-hosting. The release simplifies prompt formatting from XML to inline bracket cues and includes twelve preset voices, though users may need to accept Gemma 3 12B model licenses and provide an HF token for initial setup. AI
IMPACT Enhances creative workflows by making advanced generative audio tools more accessible on consumer hardware.
RANK_REASON New integration of an existing AI model into a popular creative workflow tool.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →