A new study published on arXiv evaluates the ability of large language models (LLMs) to determine functional equivalence between programs written in different programming languages. The research introduces a dataset called PolyHuman, comprising human-written code in C++, Java, and Python, to test LLMs' semantic understanding beyond superficial similarities. The findings indicate that current LLMs struggle with this task, exhibiting a breakdown in judgment that worsens with problem difficulty and, in some cases, showing instability or language-specific biases. AI
影响 Current LLMs do not reliably capture functional equivalence across programming languages, indicating a need for improved semantic reasoning capabilities.
排序理由 Research paper published on arXiv detailing evaluation of LLMs on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]
AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →