Kyūtai
PulseAugur coverage of Kyūtai — every cluster mentioning Kyūtai across labs, papers, and developer communities, ranked by signal.
-
Kyutai releases open-weight MuScriptor for multi-instrument music transcription
Kyutai, a French AI lab, has introduced MuScriptor, an open-weight transformer model designed for transcribing multi-instrument music into MIDI format. The model is available in three sizes, with the largest having 1.4 …
-
AI voice startup Gradium raises $100M seed, backed by Nvidia
Paris-based AI voice startup Gradium has secured $100 million in a seed funding round, with Nvidia participating as a new investor. This funding will support the company's expansion into the Bay Area to compete for tale…
-
MuScriptor: Open-weight model transcribes multi-instrument music to MIDI
Researchers have introduced MuScriptor, an open-weight model designed for multi-instrument music transcription. Unlike previous models that struggled with complex mixes or single instruments, MuScriptor can transcribe n…
-
MIRA AI model trained on Rocket League data released with demo
A new AI model called MIRA has been released, designed for multiplayer interactive world modeling and trained on data from the game Rocket League. Developed through a collaboration involving General Intuition, Kyutai, a…
-
CPU TTS benchmark: Pocket TTS shows flat RTF, UTMOS struggles with naturalness
A benchmark comparing four CPU-based text-to-speech (TTS) models—Kokoro, Supertonic, Inflect-Nano, and Pocket TTS—reveals distinct performance characteristics. Pocket TTS, utilizing a streaming language model architectu…
-
Kyutai's Pocket TTS offers CPU-based voice cloning from 5s audio
Kyutai has released Pocket TTS, a ~100M parameter streaming language model that generates audio tokens autoregressively. This model is notable for its ability to perform zero-shot voice cloning from just 5 seconds of au…
-
New open-weight MuScriptor model transcribes multi-instrument music
A new open-weight model named MuScriptor has been released for automatic music transcription, capable of converting audio recordings with multiple instruments into detailed note sequences. Developed through a collaborat…
-
Sakana AI's KAME architecture injects LLM knowledge into speech AI without latency
Sakana AI has developed KAME, a novel tandem architecture for speech-to-speech AI that aims to combine the speed of direct systems with the knowledge depth of LLM-based approaches. KAME operates with two asynchronous co…