A new study published on arXiv explores the alignment between human and large language model (LLM) evaluations of creativity. The research found that LLMs tend to focus on intrinsic qualities like novelty, aligning well with human judgments in this area. However, LLMs diverge from human evaluations when contextual information, such as social or market relevance, is required, as they are less sensitive to these factors. The study also noted that different LLMs apply distinct and varied standards, impacting their creativity judgments. AI
IMPACT LLM creativity evaluation standards may need refinement to incorporate contextual understanding for better alignment with human judgment.
RANK_REASON The cluster contains an academic paper discussing research findings on LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →