Researchers have introduced Edit2TikZ, a new benchmark designed to evaluate the capabilities of multimodal large language models (MLLMs) in editing scientific figures using TikZ code. The benchmark includes 1,548 samples covering real-world and synthetic edits, supporting both text and visual localization requests, and featuring multi-step editing with annotations. Evaluations of 14 mainstream MLLMs revealed that current models struggle with compilation success rates and preserving unrelated content, with proprietary models achieving only 75% compilation success. To address these limitations, a mixed training set called TikZEditMix and curriculum learning were developed, significantly improving the performance of compact models like Qwen3.5-4B. AI
IMPACT Highlights limitations in current MLLMs for precise code generation tasks, potentially guiding future model development for scientific visualization.
RANK_REASON The cluster describes a new benchmark and evaluation of existing models on a specific task, fitting the research category. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →