Researchers have developed CHARM, a new benchmark designed to evaluate how well large language models (LLMs) adhere to character knowledge boundaries in role-playing scenarios. The benchmark, which includes 40 characters from diverse cultural backgrounds, distinguishes between a model's awareness of out-of-scope queries and its ability to abstain from answering them. Evaluations indicate that LLMs primarily fail by providing factual, out-of-character answers despite recognizing the query is inappropriate, a phenomenon often linked to parametric overrides where the model possesses the correct information but cannot suppress it. The study also noted cultural variations in these failures, suggesting biases in how different cultural representations are embedded in model knowledge. AI
IMPACT This benchmark could lead to more robust and culturally sensitive role-playing LLMs by identifying and addressing specific failure modes.
RANK_REASON The cluster contains a research paper detailing a new benchmark for evaluating LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →