Two new research papers introduce novel approaches to source code embedding using large language models. The first, LSem2Vec, combines LLMs with sentence embedding models to extract code semantics without task-specific fine-tuning, demonstrating superior performance on various programming languages. The second paper presents jina-code-embeddings, a suite of smaller, efficient models that leverage an autoregressive backbone trained on text and code, achieving state-of-the-art results for tasks like natural language code retrieval and semantic similarity identification. AI
IMPACT These advancements could improve code analysis, retrieval, and understanding within software engineering workflows.
RANK_REASON Two academic papers published on arXiv introducing new methods for source code embedding.
- arXiv
- Han Xiao
- Hugging Face
- jina-code-embeddings
- alphaXiv
- artificial intelligence
- CatalyzeX Code Finder for Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Influence Flower
- large-language models
- LSem2Vec
- ScienceCast
- Sentence Embedding Models
- software engineering
- Source Code Embedding
- Zixiang Xian
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →