Moshi
PulseAugur coverage of Moshi — every cluster mentioning Moshi across labs, papers, and developer communities, ranked by signal.
4 day(s) with sentiment data
-
New method suppresses spurious speech in full-duplex LLMs
Researchers have identified and addressed an issue in full-duplex speech LLMs where models like Moshi and PersonaPlex inappropriately initiate speech during prolonged user silence. The problem stems from a sudden spike …
-
RetroThinker framework boosts SpeechLLM reasoning accuracy
Researchers have developed RetroThinker, a novel post-training framework designed to enhance the reasoning capabilities of speech-based large language models (SpeechLLMs). This framework enables models like Moshi to sel…
-
New AV-STE system enhances dialogue models with audio-visual speech token restoration
Researchers have developed AV-STE, a novel modular front-end system designed to enhance audio-visual speech token processing for spoken dialogue models. This system aims to improve the robustness of full-duplex dialogue…
-
New datasets and models advance full-duplex spoken dialogue research
Researchers have introduced new datasets and models focused on full-duplex spoken dialogue systems, which enable more natural, real-time conversational interactions. The DuplexDrama dataset offers over 2,000 hours of sy…
-
Mimi codec's semantic tokens linked to phonetic realizations
Researchers have investigated the Mimi codec, a component of the Moshi language model, focusing on its 2048-token semantic codebook. Their findings indicate that standard ABX experiments are insufficient for understandi…
-
Metronome system bounds AI model cache for real-time interaction stability
Researchers have developed a new system called Metronome designed to improve the real-time serving of interactive AI models. These models, such as Moshi, MiniCPM-o, and Qwen Omni, face a critical issue where sustained l…
-
New dialogue system integrates real-time facial generation with speech
Researchers have developed Moshi-Face, a novel full-duplex spoken dialogue system that integrates facial generation with audio processing. This system utilizes a VQ-VAE to encode facial data into discrete tokens and a F…
-
BayLing-Duplex enables native full-duplex speech dialogue with single LLM
Researchers have developed BayLing-Duplex, a novel full-duplex speech language model that enables simultaneous listening and speaking without relying on external turn-taking modules. This single autoregressive LLM can m…
-
Moshi dialogue models show synchronized internal states and predict turn-taking
Researchers have explored how full-duplex speech dialogue models coordinate their internal representations during interaction. By simulating dialogues between two instances of the Moshi model, they observed strong repre…
-
Thinking Machines previews interaction models for real-time AI collaboration
Thinking Machines has introduced a research preview of interaction models designed for native, real-time collaboration. These models process audio, video, and text simultaneously, allowing for continuous thought, respon…
-
New methods boost full-duplex speech models for better interaction
Researchers have developed new methods to enhance full-duplex speech models, enabling more natural and interactive conversations. One approach focuses on improving interactivity axes like pause handling and turn-taking …
-
Sakana AI's KAME architecture injects LLM knowledge into speech AI without latency
Sakana AI has developed KAME, a novel tandem architecture for speech-to-speech AI that aims to combine the speed of direct systems with the knowledge depth of LLM-based approaches. KAME operates with two asynchronous co…
-
Josh Talks launches first full-duplex Hindi conversational AI model
Researchers have developed the first open and reproducible full-duplex spoken dialogue system for the Hindi language. This system, named Human-1, adapts the Moshi architecture and was trained on over 26,000 hours of rea…