PulseAugur
EN
LIVE 22:54:12

New ORCA benchmark reveals LLMs struggle with data science code translation

A new benchmark called ORCA has been introduced to evaluate Large Language Models (LLMs) on their ability to translate code between data science libraries. The benchmark includes two settings: ORCA-MAIN with 1,600 tasks and ORCA-PROJECT with 200 tasks involving complete data science projects. Current frontier LLMs, such as Claude Opus-4.6, show limited performance, with success rates of 56.92% on ORCA-MAIN and 33.67% on ORCA-PROJECT. Researchers also proposed an intent-augmented method that improves translation accuracy by inferring source-code intent. AI

IMPACT Highlights limitations in current LLMs for code translation, potentially driving research into more capable models for data science tasks.

RANK_REASON The cluster describes a new academic paper introducing a benchmark for evaluating LLMs on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New ORCA benchmark reveals LLMs struggle with data science code translation

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Xiaolong Li, Jinyang Li, Bowen Qin, Ge Qu, Nan Huo, Xiaohan Xu, Shipei Lin, Reynold Cheng ·

    ORCA: Evaluating LLMs on Data Science Code Translation

    arXiv:2609.30749v1 Announce Type: new Abstract: Data Science Code Translation (DSCT) is the process of converting code between data science libraries while preserving functional equivalence and enabling interoperability across data science ecosystems. While Large Language Models …