speech synthesis
PulseAugur coverage of speech synthesis — every cluster mentioning speech synthesis across labs, papers, and developer communities, ranked by signal.
9 day(s) with sentiment data
-
New DoS attack targets end-to-end speech LLMs with acoustic perturbations
Researchers have developed a new denial-of-service (DoS) attack specifically targeting end-to-end (E2E) speech large language models (LLMs). Unlike previous attacks that relied on text prompt manipulation, this method i…
-
Local AI Updates: llama.cpp, PyTorch, Kimi-K3, and NVIDIA NeMo Speech 3.0
Recent updates in the local AI and open-source model space include performance enhancements for llama.cpp with CUDA fusion, addressing critical quantization bugs in PyTorch for AMD GPUs, and the trending Moonshot AI Kim…
-
AI training tools for travel industry guests and clients unveiled
The travel industry is exploring new AI-powered training technologies, with a focus on enhancing guest and client interactions. Two platforms, hospit-AI-lity and WALT, are highlighted as next-generation tools for travel…
-
AI training tools for travel industry unveiled on Mastodon
Two Mastodon posts highlight new AI-powered training technologies for the travel industry. One post introduces "hospit-AI-lity" and "WALT" as next-generation training tools for travel workers and agents, inviting feedba…
-
audio37 launches TTS and voice cloning, seeks feature ideas
The audio37 tool has been released, offering Text-to-Speech (TTS) capabilities along with voice cloning. The developer is soliciting community input on potential future features, such as speech recognition (STT) transcr…
-
Developer builds lightweight desktop GUI for local TTS models
A developer has created a lightweight desktop GUI using Tkinter to manage multiple local Text-to-Speech (TTS) engines, including Kokoro and Chatterbox. The tool allows users to split input text, clone voices from refere…
-
AI-powered meditation app generates dynamic audio using LLM and biometrics
A new meditation and sleep app called WhatsFeel utilizes AI to generate dynamic audio content in real-time, moving beyond static recordings. The app integrates LLM text generation with real-time TTS voice synthesis, ada…
-
r/LocalLLaMA users seek recommendations for ASR and TTS models
This Reddit post on r/LocalLLaMA asks for recommendations on Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) models. The original poster is currently using older models like Whisper and Kokoro with koboldcpp…
-
AssemblyAI details best practices for production voice agents
AssemblyAI has published a series of blog posts detailing best practices for building production-ready voice agents. The articles emphasize the importance of robust telemetry and diagnostic pipelines to catch regression…
-
Hugging Face launches Real World VoiceEQ benchmark for human-quality voice AI
Hugging Face has introduced Real World VoiceEQ, a new benchmark designed to evaluate the human quality of voice AI interactions. Unlike traditional benchmarks that focus on metrics like word error rate and latency, Voic…
-
Travel Tech Firm TTS Launches WALT 2.0 and AI Tools for Public Feedback
Travel Technology Solutions (TTS) has launched two AI-powered tools for public evaluation. WALT 2.0, an AI training engine for the hospitality industry, is now live and seeking user feedback. Additionally, TTS is offeri…
-
New benchmarks highlight AI speech synthesis challenges for spoofing detectors · 3 sources tracked
Researchers have developed new benchmarks to address the generalization gap in speech spoofing detection systems, which are struggling to keep pace with advanced LLM-driven text-to-speech and voice conversion technologi…
-
TTS evaluation confounded by ASR family alignment, new ensembles proposed
Researchers have identified a significant confound in evaluating text-to-speech (TTS) systems using automatic speech recognition (ASR) verifiers. The apparent quality of these verifiers is heavily influenced by the ASR …
-
Qwen-ASR-1.7B adapted for multilingual two-speaker speech recognition · 2 sources tracked
Researchers have developed a system for the MLC-SLM 2026 Challenge that adapts the Qwen3-ASR-1.7B model for multilingual, two-speaker conversational speech. The system integrates a speaker diarization front end with the…
-
WordVoice framework offers explicit, multi-dimensional control for LLM-based TTS
Researchers have introduced WordVoice, a novel framework designed to enhance control over Large Language Model (LLM)-based Text-to-Speech (TTS) systems. This system addresses the limitations of current implicit generati…
-
New framework ProPS synthesizes speaker embeddings from text prompts
Researchers have developed ProPS, a novel framework for synthesizing speaker embeddings conditioned on natural language prompts. This system converts textual descriptions of speaker profiles into sentence embeddings, wh…
-
Voice agent observability gaps hide critical audio-layer failures
Observability tools for voice agents often focus solely on the LLM component, neglecting crucial audio-layer failures. These failures, such as premature end-of-turn detection or slow barge-in detection, can cause calls …
-
New Study Explores Geometric Properties of Emotion Steering in TTS Models
Researchers have presented a novel study exploring the geometric properties of emotion control in text-to-speech (TTS) systems. The study compares speech language models (SLMs) and conditional flow-matching (CFM) module…
-
New Luxembourgish SQA system uses TTS, new expressive speech corpus released
Researchers have developed LuxSQA, a system for spoken question answering in Luxembourgish, a low-resource language. The system utilizes text-to-speech (TTS) technology to generate synthetic spoken questions, augmenting…
-
TTS evaluation shifts from naturalness to context-specific appropriateness
A new paper explores the challenges in evaluating text-to-speech (TTS) systems, moving beyond just 'naturalness' to consider 'appropriateness' within specific contexts. The research indicates that TTS systems perform we…