A new benchmark called ASAS has been developed to evaluate the safety of Arabic large language models (LLMs). The benchmark, which includes 801 human-curated prompts across eight safety categories, revealed that most tested models fail to defend against half of unsafe prompts. High-harm categories like weapons and illicit substances showed significant safety gaps, with direct and obfuscation-based attacks being the most effective. The study also found that language alignment does not easily transfer between languages and that automated safety judges perform worse than human annotators. AI
IMPACT Highlights the need for culturally specific safety evaluations for LLMs, impacting development and deployment in non-English speaking regions.
RANK_REASON Academic paper introducing a new benchmark for LLM safety evaluation. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →