A new benchmark called ORCA has been introduced to evaluate Large Language Models (LLMs) on their ability to translate code between data science libraries. The benchmark includes two settings: ORCA-MAIN with 1,600 tasks and ORCA-PROJECT with 200 tasks involving complete data science projects. Current frontier LLMs, such as Claude Opus-4.6, show limited performance, with success rates of 56.92% on ORCA-MAIN and 33.67% on ORCA-PROJECT. Researchers also proposed an intent-augmented method that improves translation accuracy by inferring source-code intent. AI
IMPACT Highlights limitations in current LLMs for code translation, potentially driving research into more capable models for data science tasks.
RANK_REASON The cluster describes a new academic paper introducing a benchmark for evaluating LLMs on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Claude Opus-4.6
- Connected Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- Litmaps
- ORCA
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →