A new benchmark called BanglaSafe has been developed to evaluate the safety of large language models (LLMs) in Bengali, the seventh most spoken language globally. The benchmark includes 879 prompts covering 17 culturally specific harm categories and five different prompting conditions. Evaluations of 18 frontier LLMs revealed that over half of the responses were unsafe or partially unsafe, with 14.7% containing strictly harmful content. Notably, the writing style within Bengali prompts had a greater impact on safety than the language switch itself, and existing safety classifiers performed poorly on Bengali content. AI
IMPACT Highlights the need for more culturally diverse safety evaluations for LLMs, potentially impacting future model development and deployment strategies.
RANK_REASON The cluster contains an academic paper introducing a new benchmark for LLM safety evaluation in a non-English language. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →