A new paper proposes a framework for measuring the creativity of frontier large language models (LLMs) in automated research settings. The proposed metrics assess both the usefulness (valueness) and novelty of LLM-generated ideas. Novelty is further broken down into exact-match, variable-level, and H-novelty, with variable-level novelty showing the strongest correlation with research performance. AI
IMPACT This research could lead to better evaluation of LLM capabilities in scientific discovery and automated research.
RANK_REASON The item is an academic paper proposing new metrics for evaluating LLM creativity. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX Code Finder for Papers
- Connected Papers
- CORE Recommender
- DagsHub
- Exact-Match P-Novelty
- Gotit.pub
- H-Novelty
- Hugging Face
- Influence Flower
- Litmaps
- ScienceCast
- scite Smart Citations
- Variable-level P-Novelty
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →