ENTITY
PromptEval
PromptEval
PulseAugur coverage of PromptEval — every cluster mentioning PromptEval across labs, papers, and developer communities, ranked by signal.
Total · 30d
0
2 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
0
1 over 90d
TIER MIX · 90D
TOPICS
RECENT · PAGE 1/1 · 2 TOTAL
-
Report: Only 8% of LLM prompts score "good"; output format is key
A report analyzing over 1,000 prompts used with large language models has revealed that only 8% achieved a "good" score (75 or higher). The most significant factor in prompt quality, contributing an average of 27 points…
-
New framework ranks AI models with statistical confidence intervals
Researchers have developed a new hierarchical framework for evaluating pretrained models on leaderboards, addressing the uncertainty and variability in performance across different tasks. This method constructs statisti…