PulseAugur
EN
LIVE 21:01:01

Report: Only 8% of LLM prompts score "good"; output format is key

A report analyzing over 1,000 prompts used with large language models has revealed that only 8% achieved a "good" score (75 or higher). The most significant factor in prompt quality, contributing an average of 27 points, is clearly defining the output format. Robustness emerged as the weakest dimension across most prompts, with 9 out of 10 failing to handle ambiguous or unexpected input effectively. Prompts that are too short or rely on vague descriptors like "engaging" tend to perform poorly, while fields with higher stakes, such as healthcare, produce more carefully constructed prompts. AI

IMPACT Highlights critical areas for improving LLM prompt engineering, suggesting a focus on output formatting and robustness for better results.

RANK_REASON Analysis of prompt quality based on a dataset of scored prompts. [lever_c_demoted from research: ic=1 ai=1.0]

Read on dev.to — LLM tag →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Report: Only 8% of LLM prompts score "good"; output format is key

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
Analysis of prompt quality based on a dataset of scored prompts. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
product, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
80 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [1]

  1. dev.to — LLM tag TIER_1 English(EN) · Francisco Ferreira ·

    The Prompt Quality Report: What 1,000 Scored Prompts Reveal

    <blockquote> <p><strong>Quick answer:</strong> The PromptEval Prompt Quality Report scored over 1,000 real prompts across 12 use cases. The average was 52 out of 100, and only 8% reached "good" (75+). The strongest single predictor of a good prompt is whether it defines its outpu…