A new benchmark called WISRD has been developed to test the rule-discovery capabilities of image-editing models. The benchmark includes 11 core tasks and supplementary probes designed to assess spatial manipulation, pattern reasoning, and logical inference. In evaluations, Nano Banana Pro outperformed other models, achieving a 48.7% pass rate on a subset of the benchmark, significantly higher than Qwen Image Edit and FLUX.2 Klein 4B. AI
IMPACT This benchmark could drive improvements in AI's ability to understand and manipulate visual information, potentially leading to more sophisticated image editing and reasoning tools.
RANK_REASON The item describes a new academic benchmark and evaluation of AI models on that benchmark. [lever_c_demoted from research: ic=1 ai=1.0]
- FLUX.2 Klein 4B API
- FLUX.2 Klein 4B open-weight
- InstructPix2Pix
- Nano Banana Pro
- Qwen Image Edit
- RAVEN
- WISRD
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →