PulseAugur
EN
LIVE 18:52:15

AssemblyAI integrates Universal-3.5 Pro Realtime with Agora for voice agents

AssemblyAI has released a guide detailing how to integrate its Universal-3.5 Pro Realtime speech-to-text model with Agora's real-time audio transport platform. This integration allows developers to add low-latency, speaker-aware transcription to Agora calls without modifying client-side code. The process involves a Python server bot that joins an Agora channel, captures raw audio, and streams it to AssemblyAI's WebSocket endpoint for transcription. AssemblyAI's model reportedly achieves a 6.99% Word Error Rate (WER) on real voice-agent conversations, outperforming competitors like Deepgram Flux, ElevenLabs Scribe v2, and Google Chirp3 in benchmarks. AI

IMPACT Enables developers to easily add advanced speech-to-text capabilities to real-time communication platforms.

RANK_REASON Blog post detailing integration of existing products, not a new product or model release.

Read on AssemblyAI blog →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

AssemblyAI integrates Universal-3.5 Pro Realtime with Agora for voice agents

COVERAGE [1]

  1. AssemblyAI blog TIER_1 English(EN) ·

    How to Build an Agora Voice Agent with AssemblyAI

    Add speaker-aware, low-latency transcription to Agora calls—stream raw PCM to AssemblyAI's Universal-3.5 Pro Realtime with a server bot, no client code changes.