PulseAugur
实时 03:42:58
English(EN) The Human Creativity Benchmark

新的人类创造力基准用于评估AI的品味和一致性

提出了一项名为“人类创造力基准”(HCB)的新基准,用于评估AI在创意领域的模型。与将评估者分歧视为噪音的传统基准不同,HCB同时保留了收敛性(专业人士达成一致的方面)和发散性(个人品味不同的方面)。该基准收集了来自领域专业人士在各种创意任务和工作流程阶段的判断,旨在通过区分需要技术正确性和允许主观可控性的领域,为AI性能提供更具可操作性的见解。 AI

影响 该基准可能导致对创意AI进行更细致的评估,从而可能指导开发出在技术熟练度和主观美学品质之间取得更好平衡的模型。

排序理由 该集群包含一篇介绍AI评估新基准的学术论文。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 5 个来源。 我们如何撰写摘要 →

新的人类创造力基准用于评估AI的品味和一致性

报道来源 [5]

  1. arXiv cs.AI TIER_1 English(EN) · Zhaoyi Joey Hou, Bowei Alvin Zhang, Yining Lu, Bhiman Kumar Baghel, Anneliese Brei, Ximing Lu, Meng Jiang, Faeze Brahman, Snigdha Chaturvedi, Haw-Shiuan Chang, Daniel Khashabi, Xiang Lorraine Li ·

    CreativityPrism:大型语言模型创造力的跨领域评估框架

    arXiv:2510.20091v3 Announce Type: replace-cross Abstract: Creativity is often seen as a hallmark of human intelligence. While large language models(LLMs) are increasingly perceived as generating creative text, there is still no cross-domain and scalable framework to evaluate thei…

  2. arXiv cs.CL TIER_1 English(EN) · Roger Beaty, Vijeta Deshpande, Clin K. Y. Lai, Anna Attuch, Namrata Shivagunde, Swastik Roy, Rajkumar Pujari, Paul V. DiStefano, Sherin Muckatira, Claire E. Stevenson, Mikhail Gronas, Anna Rumshisky ·

    AGC-Bench:衡量通用人工智能创造力

    arXiv:2607.01152v1 Announce Type: new Abstract: Creativity research has debated whether creativity is domain-specific (e.g., visual, writing, science), and if it is psychometrically separable from general intelligence. Both questions now apply to LLMs, but a unified benchmark of …

  3. arXiv cs.CL TIER_1 English(EN) · Anna Rumshisky ·

    AGC-Bench:衡量通用人工智能创造力

    Creativity research has debated whether creativity is domain-specific (e.g., visual, writing, science), and if it is psychometrically separable from general intelligence. Both questions now apply to LLMs, but a unified benchmark of AI creativity remains elusive. We introduce AGC-…

  4. arXiv cs.AI TIER_1 English(EN) · Aspen Hopkins, Allison Nulty, Alexandria Minetti, Anoop Pakki, Angad Singh ·

    人类创造力基准

    arXiv:2606.30561v1 Announce Type: new Abstract: Modern AI evaluation frameworks treat evaluator disagreement as noise to be resolved. In creative domains, professional disagreement reflects genuine differences in taste, not measurement error. We argue that evaluating creative AI …

  5. arXiv cs.AI TIER_1 English(EN) · Angad Singh ·

    人类创造力基准

    Modern AI evaluation frameworks treat evaluator disagreement as noise to be resolved. In creative domains, professional disagreement reflects genuine differences in taste, not measurement error. We argue that evaluating creative AI requires preserving two distinct signals: conver…