Researchers have introduced MobileForge, a new benchmark designed to evaluate multimodal large language models' ability to generate complete, functional mobile applications from visual designs. Existing benchmarks fall short by focusing on single pages and failing to assess cross-page navigation or project-wide maintainability. MobileForge addresses these limitations by including real mobile apps, human-reviewed screens, navigation specifications, and a five-axis evaluation framework covering buildability, navigation, visual fidelity, code maintainability, and efficiency. Initial tests on six frontier multimodal LLMs show that while current models can generate compilable app projects that reach the correct pages, interactive navigation remains unreliable, and visual fidelity and maintainability still require significant improvement. AI
IMPACT This benchmark will drive improvements in multimodal LLMs for software development, potentially accelerating the creation of complex mobile applications.
RANK_REASON The cluster contains an academic paper introducing a new benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →