Two new research papers propose standardized protocols for evaluating real-time voice agents. The first paper, "Evaluating Real-Time Voice Agents: From Component Quality to Grounded Outcomes," introduces the TRG (Timing-Recovery-Grounded) standard, which measures an agent by its timing, post-disruption recovery, and state-verified outcomes, emphasizing a shift from component quality to grounded results. The second paper, "Inquesto Score: A reliability Protocol For Voice Agents," presents the Inquesto Score (IS), a protocol focused on measuring voice-agent reliability as the percentage of calls that achieve the caller's goal without functional failure, highlighting the need for evidence beyond transcripts and explicit treatment of deployment conditions. AI
IMPACT Standardized evaluation protocols are crucial for advancing the reliability and deployment of voice agents in real-world applications.
RANK_REASON Two academic papers proposing new evaluation protocols for voice agents.
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- DagsHub
- Gotit.pub
- Hugging Face
- Inquesto Score
- Litmaps
- ScienceCast
- scite Smart Citations
- TRG
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →