AssemblyAI hosted a meetup in San Francisco focused on building practical voice agents, featuring insights from production teams at Retell and Super. Key takeaways included the preference for cascading architectures for flexibility and customization, the critical importance of maintaining a 1-1.5 second latency budget for natural conversation flow, and the layered approach to model evaluation. This evaluation process spans automated metrics for speech-to-text, curated hard cases and LLM-as-judge for language models, and ultimately human "vibe checks" for text-to-speech quality. AI
IMPACT Provides practical insights for developers building voice agents, focusing on architectural choices, latency management, and evaluation strategies.
RANK_REASON Blog post summarizing takeaways from a company-hosted meetup, not a primary release or significant industry event.
- Adam Schuld
- AssemblyAI
- Context Carryover
- Retell
- Ryan Seams
- Super
- Universal-3.5 Pro Realtime
- Zhongren Shao
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →