Jieyue (Step) has launched its new generation of large voice models, the StepAudio 3 series, featuring five distinct models: StepAudio 3 Realtime, StepAudio 3 ASR, StepAudio 3 TTS, StepAudio 3 Gen, and StepAudio 3 Music. These models cover a wide range of applications including real-time voice interaction, speech recognition, realistic voice generation, and music creation. Notably, several of these models have achieved top global rankings on the Artificial Analysis benchmark, demonstrating advanced capabilities in areas like conversational dynamics and speech reasoning. AI
IMPACT Sets new benchmarks for voice AI capabilities, potentially accelerating adoption in real-time interaction and content creation.
RANK_REASON New model release from a significant AI lab with benchmark claims. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
- Artificial Analysis
- Conversational Dynamics
- Speech Reasoning
- StepAudio 3
- StepAudio 3 ASR
- StepAudio 3 Gen
- StepAudio 3 Music
- StepAudio 3 Realtime
- StepAudio 3 TTS
- Zhang Haipeng
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →