Researchers have developed SpIDER, a novel dense retrieval method that enhances the ability of large language models (LLMs) to locate relevant code segments within large codebases. Unlike existing methods that focus solely on semantic similarity, SpIDER integrates graph-based exploration of codebase structures, such as containment and call relationships. This approach is validated by SpIDER-Bench, a new benchmark dataset designed for graph-structured code retrieval across multiple programming languages including Python, Java, JavaScript, and TypeScript. Empirical results demonstrate significant improvements in retrieval accuracy, particularly in Recall@20, by leveraging the spatial and structural information within the code. AI
IMPACT Enhances LLM capabilities in code understanding and retrieval, potentially improving developer productivity and agent performance.
RANK_REASON Research paper introducing a new method and benchmark for software issue localization. [lever_c_demoted from research: ic=1 ai=1.0]
- BM25
- Java
- Javascript
- large language model
- Multi-SWE-bench
- Python
- Shravan Sunil Chaudhari
- SpIDER-Bench
- SWEBench-Verified
- SWEPolyBench
- TypeScript
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →