ENTITY
AD2-Bench
AD2-Bench
PulseAugur coverage of AD2-Bench — every cluster mentioning AD2-Bench across labs, papers, and developer communities, ranked by signal.
Total · 30d
2
2 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
2
2 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D
2 day(s) with sentiment data
RECENT · PAGE 1/1 · 2 TOTAL
-
New benchmark AD2-Bench evaluates MLLM trustworthiness in complex urban scenes
Researchers have introduced AD2-Bench, a new evaluation benchmark designed to assess the trustworthiness of Multimodal Large Language Models (MLLMs) in complex urban environments. Unlike existing benchmarks that only ev…
-
New benchmarks and methods enhance multimodal AI reasoning and trustworthiness · 4 sources tracked
Researchers are developing new methods to improve the reliability and trustworthiness of multimodal large language models (MLLMs). One approach, VERDICT, uses disagreement among multiple verifiers to identify errors in …