Researchers have introduced SciCodePile, a substantial 128GB corpus of scientific code, to address the limitations of existing datasets in evaluating large language models' (LLMs) ability to generate scientific code. This corpus, derived from over 37,000 public repositories, is accompanied by an executable benchmark of 200 tasks designed for functional verification. Evaluations of 15 LLMs revealed significant challenges, with the best models achieving only a 38.13% CodeBLEU score on completion tasks and a 12.30% Pass@1 rate on executable code generation, indicating a considerable gap between current LLMs and reliable scientific code generation. AI
IMPACT Highlights the current limitations of LLMs in generating complex scientific code, suggesting areas for future model development and training.
RANK_REASON The cluster describes a new academic paper introducing a dataset and benchmark for evaluating LLMs on scientific code generation. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →