A new benchmark called FigCodeBench has been developed to evaluate the capabilities of Multimodal Large Language Models (MLLMs) in reproducing complex visual figures and generating corresponding code. This framework addresses the gap in current benchmarks by integrating visual understanding and code generation, moving beyond isolated assessments. Experiments conducted on 24 proprietary and open-source MLLMs, including Gemini-3.1 Pro and GPT-5.4, revealed a significant performance drop across various programming languages and difficulty levels, offering insights into their limitations. AI
IMPACT Highlights limitations in current MLLMs for complex visual-to-code tasks, potentially guiding future model development.
RANK_REASON New academic paper introducing a novel benchmark and evaluation framework for MLLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- DagsHub
- FigCodeBench
- Gemini-3.1 Pro
- Gotit.pub
- GPT-5.4
- Hugging Face
- Kimi K2.5
- Litmaps
- MLLMs
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →