Researchers have introduced ProcArena, a new benchmark designed to evaluate large language models (LLMs) on their ability to generate PL/SQL code from natural language. The benchmark covers both direct generation and interactive development scenarios across PostgreSQL and Oracle databases, encompassing tasks like initial development, code modification, debugging, and optimization. Initial evaluations of seven LLMs revealed that the best-performing models achieved only 62.2% and 57.8% accuracy in direct and interactive modes, respectively, highlighting the significant challenges that remain in realistic NL-to-PL/SQL development. AI
IMPACT Highlights current limitations of LLMs in complex code generation tasks, suggesting areas for future research and development.
RANK_REASON The cluster contains a research paper introducing a new benchmark for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- Iterative Logic Enhancement
- Knowledge integration
- LLMs
- natural language
- Oracle
- PL/SQL
- PostgreSQL
- ProcArena
- Requirement Perturbation
- Solver-User Simulator
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →