PulseAugur
实时 05:35:13

LLMs fail to reliably detect cross-language code equivalence, study finds

A new study published on arXiv evaluates the ability of large language models (LLMs) to determine functional equivalence between programs written in different programming languages. The research introduces a dataset called PolyHuman, comprising human-written code in C++, Java, and Python, to test LLMs' semantic understanding beyond superficial similarities. The findings indicate that current LLMs struggle with this task, exhibiting a breakdown in judgment that worsens with problem difficulty and, in some cases, showing instability or language-specific biases. AI

影响 Current LLMs do not reliably capture functional equivalence across programming languages, indicating a need for improved semantic reasoning capabilities.

排序理由 Research paper published on arXiv detailing evaluation of LLMs on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

LLMs fail to reliably detect cross-language code equivalence, study finds

本文如何被排名

Signal score
43 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
Research paper published on arXiv detailing evaluation of LLMs on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.AI TIER_1 (CA) · Hui Sun, Anderson Uch\^oa, Rohit Gheyi, Wesley K. G. Assun\c{c}\~ao ·

    跨语言代码功能等价性语言模型评估

    arXiv:2608.23961v1 Announce Type: cross Abstract: Background: Large Language Models (LLMs) have demonstrated strong performance across a variety of code-understanding tasks, leading many to believe that they can reason about program semantics. However, existing evaluations primar…