A new survey paper published on arXiv examines the capabilities of large language models (LLMs) in software engineering and security. The paper highlights that while LLMs are advancing towards repository-scale agents capable of complex tasks, their evaluation remains fragmented. Current assessments often focus on either functional task completion in software engineering or vulnerability detection in software security, without adequately bridging the gap between the two. The survey identifies common validity threats in existing research and proposes a minimum reporting protocol to improve cross-study comparability, advocating for jointly secure-and-functional benchmarks and reproducible agent evaluation. AI
IMPACT Highlights the need for unified benchmarks to assess LLM capabilities in both functional and security aspects of software development.
RANK_REASON The item is a survey paper published on arXiv detailing research findings and proposing a future research agenda. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX Code Finder for Papers
- Connected Papers
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- large-language models
- Litmaps
- ScienceCast
- scite Smart Citations
- software engineering
- software security
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →