PulseAugur
中
实时 00:00:42

深度学习模型在大规模代码检索方面存在困难,新论文发现

一篇题为“Recall Before Rerank”的新研究论文评估了深度学习模型在大规模代码到代码检索方面的性能。该研究强调了当前模型在处理跨多种编程语言的TB级源代码集合时,在精度和可扩展性方面存在的局限性。研究人员提出了基于LLM的代码规范化和查询重写技术,以提高效果较差模型的性能,并质疑资源受限环境下专门代码LLM的可行性。 AI

影响 强调了当前LLM在代码检索方面的局限性,为开发人员在可扩展性和精度方面指出了改进方向。

排序理由 该集群包含一篇在arXiv上发表的研究论文。

在 arXiv cs.IR (Information Retrieval) 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

深度学习模型在大规模代码检索方面存在困难,新论文发现

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群包含一篇在arXiv上发表的研究论文。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
106 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [2]

  1. arXiv cs.CL TIER_1 English(EN) · Leonardo Venuta, Francesco Tosoni, Paolo Ferragina ·

    召回优先于重排:大规模代码到代码检索的深度学习模型基准测试

    arXiv:2606.27401v1 Announce Type: cross Abstract: Semantic code search and clone detection are essential for software development, maintenance, and reuse. This paper evaluates the effectiveness, efficiency, and scalability of contemporary deep learning models for first-stage reca…

  2. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Paolo Ferragina ·

    召回优先于重排:大规模代码到代码检索的深度学习模型基准测试

    Semantic code search and clone detection are essential for software development, maintenance, and reuse. This paper evaluates the effectiveness, efficiency, and scalability of contemporary deep learning models for first-stage recall in large-scale code-to-code search engines. Ben…