A new study published on arXiv evaluates the ability of large language models (LLMs) to determine functional equivalence between programs written in different programming languages. The research introduces a dataset called PolyHuman, comprising human-written code in C++, Java, and Python, to test LLMs' semantic understanding beyond superficial similarities. The findings indicate that current LLMs struggle with this task, exhibiting a breakdown in judgment that worsens with problem difficulty and, in some cases, showing instability or language-specific biases. AI
IMPACT Current LLMs do not reliably capture functional equivalence across programming languages, indicating a need for improved semantic reasoning capabilities.
RANK_REASON Research paper published on arXiv detailing evaluation of LLMs on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →