Researchers have developed RAG-PIBench, a new benchmark designed to evaluate the effectiveness of prompt-injection detection in retrieval-augmented generation (RAG) systems. The benchmark includes 4,876 contextual examples across distinct training, validation, and testing sets, employing a leakage-aware construction process. Evaluations showed that the DistilBERT model achieved the highest performance, with a F1 score of 0.896 and a PR-AUC of 0.968, though traditional methods like TF-IDF SVM and logistic regression also demonstrated competitive results. AI
IMPACT This benchmark will help improve the security of RAG systems against prompt injection attacks.
RANK_REASON The item is an academic paper detailing a new benchmark for AI security. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- DistilBERT
- Hugging Face
- logistic regression model
- RAG-PIBench
- retrieval-augmented generation
- TF-IDF SVM
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →