PulseAugur
EN
LIVE 11:27:55

New Human Creativity Benchmark for AI Evaluates Taste and Agreement

A new benchmark called the Human Creativity Benchmark (HCB) has been proposed to evaluate AI models in creative domains. Unlike traditional benchmarks that treat evaluator disagreement as noise, HCB preserves both convergence (where professionals agree) and divergence (where individual taste varies). The benchmark collects judgments from domain professionals across various creative tasks and workflow phases, aiming to provide more actionable insights into AI performance by distinguishing between areas requiring technical correctness and those allowing for subjective steerability. AI

IMPACT This benchmark could lead to more nuanced evaluations of creative AI, potentially guiding development towards models that better balance technical proficiency with subjective aesthetic qualities.

RANK_REASON The cluster contains an academic paper introducing a new benchmark for AI evaluation.

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 5 sources. How we write summaries →

New Human Creativity Benchmark for AI Evaluates Taste and Agreement

COVERAGE [5]

  1. arXiv cs.AI TIER_1 English(EN) · Zhaoyi Joey Hou, Bowei Alvin Zhang, Yining Lu, Bhiman Kumar Baghel, Anneliese Brei, Ximing Lu, Meng Jiang, Faeze Brahman, Snigdha Chaturvedi, Haw-Shiuan Chang, Daniel Khashabi, Xiang Lorraine Li ·

    CreativityPrism: A Cross-Domain Evaluation Framework for Large Language Model Creativity

    arXiv:2510.20091v3 Announce Type: replace-cross Abstract: Creativity is often seen as a hallmark of human intelligence. While large language models(LLMs) are increasingly perceived as generating creative text, there is still no cross-domain and scalable framework to evaluate thei…

  2. arXiv cs.CL TIER_1 English(EN) · Roger Beaty, Vijeta Deshpande, Clin K. Y. Lai, Anna Attuch, Namrata Shivagunde, Swastik Roy, Rajkumar Pujari, Paul V. DiStefano, Sherin Muckatira, Claire E. Stevenson, Mikhail Gronas, Anna Rumshisky ·

    AGC-Bench: Measuring Artificial General Creativity

    arXiv:2607.01152v1 Announce Type: new Abstract: Creativity research has debated whether creativity is domain-specific (e.g., visual, writing, science), and if it is psychometrically separable from general intelligence. Both questions now apply to LLMs, but a unified benchmark of …

  3. arXiv cs.CL TIER_1 English(EN) · Anna Rumshisky ·

    AGC-Bench: Measuring Artificial General Creativity

    Creativity research has debated whether creativity is domain-specific (e.g., visual, writing, science), and if it is psychometrically separable from general intelligence. Both questions now apply to LLMs, but a unified benchmark of AI creativity remains elusive. We introduce AGC-…

  4. arXiv cs.AI TIER_1 English(EN) · Aspen Hopkins, Allison Nulty, Alexandria Minetti, Anoop Pakki, Angad Singh ·

    The Human Creativity Benchmark

    arXiv:2606.30561v1 Announce Type: new Abstract: Modern AI evaluation frameworks treat evaluator disagreement as noise to be resolved. In creative domains, professional disagreement reflects genuine differences in taste, not measurement error. We argue that evaluating creative AI …

  5. arXiv cs.AI TIER_1 English(EN) · Angad Singh ·

    The Human Creativity Benchmark

    Modern AI evaluation frameworks treat evaluator disagreement as noise to be resolved. In creative domains, professional disagreement reflects genuine differences in taste, not measurement error. We argue that evaluating creative AI requires preserving two distinct signals: conver…