PulseAugur
EN
LIVE 07:22:33
ENTITY AD2-Bench

AD2-Bench

PulseAugur coverage of AD2-Bench — every cluster mentioning AD2-Bench across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
2
2 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
2
2 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

2 day(s) with sentiment data

RECENT · PAGE 1/1 · 2 TOTAL
  1. TOOL · CL_228991 ·

    New benchmark AD2-Bench evaluates MLLM trustworthiness in complex urban scenes

    Researchers have introduced AD2-Bench, a new evaluation benchmark designed to assess the trustworthiness of Multimodal Large Language Models (MLLMs) in complex urban environments. Unlike existing benchmarks that only ev…

  2. RESEARCH · CL_193326 ·

    New benchmarks and methods enhance multimodal AI reasoning and trustworthiness · 4 sources tracked

    Researchers are developing new methods to improve the reliability and trustworthiness of multimodal large language models (MLLMs). One approach, VERDICT, uses disagreement among multiple verifiers to identify errors in …