A new benchmark called CS-Guard has been developed to evaluate the effectiveness of Large Language Model (LLM) guardrails in preventing the generation of malicious code. The benchmark includes over 1000 prompts for text-to-code generation and 331 prompts for code-to-code generation, incorporating jailbreak attacks and a novel fictional scenario attack. Empirical testing of nine guardrails across seven LLMs revealed significant vulnerabilities, with attack success rates reaching up to 50% for text-to-code and nearly 100% for code-to-code generation, raising concerns for real-world software development. AI
IMPACT Highlights critical security vulnerabilities in LLMs, potentially impacting the adoption of AI in sensitive code generation tasks.
RANK_REASON The item describes a new academic benchmark and research paper evaluating LLM security guardrails. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- CatalyzeX Code Finder for Papers
- code generation
- Connected Papers
- CORE Recommender
- CS-Guard
- DagsHub
- Financial Services Agency
- Gotit.pub
- Hugging Face
- Large language models
- Litmaps
- LLM guardrails
- malware
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →