PulseAugur
中
实时 14:58:57
English(EN) ORCA: Evaluating LLMs on Data Science Code Translation

新基准揭示LLM在代码生成和翻译方面的局限性

引入了两个新的基准测试E2E-SWE和ORCA,以评估大型语言模型(LLM)在复杂编码任务中的能力。E2E-SWE侧重于LLM从头开始生成整个软件存储库的能力,使用11种编程语言对13个前沿模型进行了测试,结果显示性能差异显著。另一方面,ORCA评估LLM在数据科学库之间翻译代码的熟练程度,结果表明即使是像Claude Opus-4.6这样的先进模型也难以胜任这项任务,这促使开发了一种意图增强的翻译方法。 AI

影响 这些基准测试突显了LLM在复杂软件工程任务中当前的局限性,指明了AI编码能力未来研究和发展的方向。

排序理由 该集群包含两篇介绍用于评估LLM的新基准的学术论文。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新基准揭示LLM在代码生成和翻译方面的局限性

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群包含两篇介绍用于评估LLM的新基准的学术论文。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
9 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [2]

  1. arXiv cs.AI TIER_1 English(EN) · Hantian Ding, Chloe Bi, Jiacheng Zhu, John Yang, Matt Deitke, Pengcheng Yin, Zijian Wang, Rui Hou ·

    E2E-SWE:从零开始构建可运行代码库的 LLM 基准测试

    arXiv:2609.38335v1 Announce Type: cross Abstract: Coding agents powered by large language models (LLMs) are evolving from making localized code changes to developing complete software repositories. However, evaluating repository-scale generation remains challenging: tasks must de…

  2. arXiv cs.AI TIER_1 English(EN) · Xiaolong Li, Jinyang Li, Bowen Qin, Ge Qu, Nan Huo, Xiaohan Xu, Shipei Lin, Reynold Cheng ·

    ORCA:在数据科学代码翻译方面评估LLM

    arXiv:2609.30749v1 Announce Type: new Abstract: Data Science Code Translation (DSCT) is the process of converting code between data science libraries while preserving functional equivalence and enabling interoperability across data science ecosystems. While Large Language Models …