NVIDIA has launched NemotronLabs VoiceChat 11B, an open-source, full-duplex speech-to-speech model designed for real-time conversational AI. This unified model integrates speech recognition, language understanding, and speech synthesis, achieving a turn-taking latency of approximately 448 ms. A key feature is its ability to perform tool calling concurrently with ongoing conversation, allowing for seamless integration of external functions without interrupting the dialogue. AI
IMPACT Enables more natural and responsive voice interactions by reducing latency and allowing concurrent tool use in conversational AI applications.
RANK_REASON NVIDIA released a new open-source model with detailed technical specifications and performance metrics. [lever_c_demoted from frontier_release: ic=2 ai=1.0]
Read on Mastodon — fosstodon.org →
- A100
- Audio Flamingo 3
- Full-Duplex-Bench 1.0
- NemotronLabs VoiceChat 11B
- Nemotron-Speech-Streaming-En-0.6b
- NVIDIA
- Nvidia B200
- NVIDIA H100
- NVIDIA Nemotron Nano v2
- RTX 6000 Pro
- SALM-Duplex
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →