A new benchmark called CLI-Tool-Bench has been developed to evaluate the capability of Large Language Models (LLMs) in generating complete, runnable software systems from scratch, addressing the limitations of existing benchmarks that rely on predefined structures and rigid testing methods. The benchmark features 94 real-world repositories and uses a black-box differential testing framework to assess end-to-end user-centric behavior. Initial evaluations of seven state-of-the-art LLMs showed a maximum success rate of only 43.8%, indicating that autonomous 0-to-1 software generation remains a significant challenge, with models tending to produce monolithic code structures and not necessarily improving performance with higher token consumption. AI
IMPACT Highlights significant limitations in current LLMs for autonomous software development, suggesting a need for improved architectures and training methodologies.
RANK_REASON The cluster contains an academic paper detailing a new benchmark and evaluation of LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →