PulseAugur
EN
LIVE 23:35:36
ENTITY EVA-Bench

EVA-Bench

PulseAugur coverage of EVA-Bench — every cluster mentioning EVA-Bench across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
1
1 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
0
0 over 90d
TIER MIX · 90D
TOPICS
TIMELINE
  1. 2026-05-13 research_milestone Researchers released EVA-Bench, a new end-to-end framework for evaluating voice agents. source
SENTIMENT · 30D

1 day(s) with sentiment data

RECENT · PAGE 1/1 · 5 TOTAL
  1. FRONTIER RELEASE · CL_255943 ·

    Google launches Gemini 3.8 Live and Extended Thinking audio models

    Google has launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, new advanced audio models designed to enhance voice agent capabilities and create more natural AI conversations. Gemini 3.8 Live focuses on scal…

  2. SIGNIFICANT · CL_83825 ·

    xAI launches Grok Voice with human-like performance and lower cost

    xAI has released Grok Voice, a new voice AI model that boasts state-of-the-art performance with human-like qualities. The model achieves top-tier accuracy and user experience simultaneously, outperforming competitors on…

  3. RESEARCH · CL_71082 ·

    Hugging Face expands voice agent benchmark to 3 domains, 121 tools

    Hugging Face has released EVA-Bench Data 2.0, an expanded benchmark for evaluating voice agents. This new version broadens its scope to three enterprise domains: Airline Customer Service Management, Enterprise IT Servic…

  4. TOOL · CL_30709 ·

    New EVA-Bench framework evaluates voice agent performance

    Researchers have introduced EVA-Bench, a new framework designed to comprehensively evaluate voice agents. This system addresses key challenges by generating realistic simulated conversations and measuring quality across…

  5. RESEARCH · CL_44365 ·

    New benchmarks and platforms advance voice agent evaluation and development

    New research introduces EVA-Bench, a comprehensive framework for evaluating voice agents, addressing challenges in simulating realistic conversations and measuring performance across various failure modes. Simultaneousl…