Researchers have introduced KaliBench, a new benchmark designed to evaluate the ability of large language models (LLMs) to generate precise command-line interface (CLI) commands for cybersecurity tools. The benchmark includes 8,504 query-command pairs across 1,642 tools, focusing on accurate tool selection and argument construction. Current open-weight models struggle with this task, achieving less than 42% exact-command accuracy. However, fine-tuning with KaliBench's verifiable rewards significantly improved an 8B model's performance. AI
IMPACT This benchmark could accelerate the development of more capable LLMs for cybersecurity operations by providing a standardized evaluation for CLI command generation.
RANK_REASON The cluster contains an academic paper introducing a new benchmark for evaluating LLM performance on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →