A new benchmark, extending CruxEval, has been developed to evaluate AI models' ability to predict the final output and internal state of real-world programs. This benchmark comprises 400 cases from 371 Python and C++ programs, tested across various model families. Results indicate that models with reasoning capabilities significantly outperform those without, achieving high accuracy on final output prediction and demonstrating the benchmark's effectiveness in exposing errors. AI
IMPACT This benchmark could lead to more robust AI models capable of understanding and predicting program execution, improving code analysis and generation tools.
RANK_REASON The cluster describes a new academic paper introducing a benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX Code Finder for Papers
- CORE Recommender
- CPP
- CruxEval
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- Python
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →