PulseAugur
实时 11:01:56
English(EN) Benchmarking Frontier Text-to-Image Models on Image-Description Prompts

Gemini 3 Pro Image 在文本到图像基准测试中领先,击败 FLUX.2

一篇新论文在具有挑战性的图像描述提示上对四种领先的文本到图像模型——Hunyuan 3.0Gemini 3 Pro ImageBlack Forest Labs FLUX.2Ideogram 3.0——进行了基准测试。使用来自 Sample Dataset (DSD) 的 48 个复杂提示进行的评估显示,Gemini 3 Pro Image 以 84.8/100 的得分成为最佳表现者,紧随其后的是 FLUX.2,得分为 82.3/100。研究确定物体计数错误和几何伪影是顶级模型的首要失败点,而 Ideogram 3.0 和 Hunyuan 3.0 在文本乱码和遗漏元素方面则面临更多困难。 AI

影响 为评估文本到图像模型中复杂提示的遵循情况建立了一个基准,并突出了未来发展的方向。

排序理由 该集群是一篇评估现有模型在基准测试中表现的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Gemini 3 Pro Image 在文本到图像基准测试中领先,击败 FLUX.2

报道来源 [1]

  1. arXiv cs.CV TIER_1 English(EN) · Sajjad Abdoli, Ghassan Al-Sumaidaee, Ahmed Rashad ·

    在图像描述提示上对前沿文本到图像模型进行基准测试

    arXiv:2608.14976v1 Announce Type: new Abstract: Text-to-image models are typically reported on average-case prompts, which understates the gap between systems on compositionally demanding requests involving precise object counts, multi-object attribute binding, legible embedded t…