ENTITY
BenchMIRT
BenchMIRT
PulseAugur coverage of BenchMIRT — every cluster mentioning BenchMIRT across labs, papers, and developer communities, ranked by signal.
Total · 30d
2
2 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
1
1 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D
1 day(s) with sentiment data
RECENT · PAGE 1/1 · 2 TOTAL
-
BenchMIRT tool questions LLM safety evaluations, finds reasoning bias
A new tool called BenchMIRT has been developed to evaluate the effectiveness of LLM safety and capability assessments. Initial testing on the barbecue (BBQ) social bias evaluation revealed that the questions primarily d…
-
New BenchMIRT benchmark questions LLM evaluation methods
A new benchmark called BenchMIRT, developed by the Allen Institute for Artificial Intelligence and presented by Hugging Face, aims to re-evaluate how Large Language Models (LLMs) are assessed. The benchmark questions th…