CosyVoice2
PulseAugur coverage of CosyVoice2 — every cluster mentioning CosyVoice2 across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New X2-NativeCursor system improves text-to-speech progress tracking
Researchers have developed X2-NativeCursor, a novel system for tracking text progress in incremental text-to-speech (TTS) applications. This lightweight observer operates by analyzing native speech tokens before wavefor…
-
New attack reveals severe privacy risks in fine-tuned TTS models
Researchers have developed a new black-box membership inference attack (MIA) framework specifically designed for fine-tuned Text-to-Speech (TTS) models. This framework addresses challenges in query generation and repres…
-
NouveauVoice framework enhances voice anonymization with diverse pseudo-speakers
Researchers have developed NouveauVoice, a new framework designed to generate diverse pseudo-speakers for voice anonymization. This system utilizes a Hierarchical Deep Variational Autoencoder (NVAE) and can be integrate…
-
New TTS method boosts emotion control accuracy by 12%
Researchers have developed a new method called Cross-modal Consistency Guided Classifier-Free Guidance (CCG-CFG) to improve emotion control in auto-regressive Text-to-Speech (TTS) models. This technique dynamically adju…
-
Consumer-grade graphics cards can quickly get started! MiniCPM-o 4.5 from Mianbi Intelligent releases technical report
MiniCPM-o 4.5 is a new 9B parameter omni-modal large language model designed for real-time, full-duplex interaction. It can simultaneously process and generate audio, video, and text, enabling proactive behaviors and co…