Researchers have compiled a dataset of 57.5K unique transactional prompts from GitHub, focusing on reproducible natural language instructions integrated into software. They developed a structured ontology to analyze these prompts, revealing diverse usage patterns across languages, domains, tasks, and modalities, with a typical Zipf-like distribution. The dataset and an exploration interface are being released to facilitate empirical study of prompts, with a comprehensive error analysis conducted to ensure annotation quality. AI
IMPACT This dataset could enable deeper understanding and optimization of prompt engineering for LLMs.
RANK_REASON The item is an academic paper detailing a new dataset and methodology. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →