Researchers have introduced KoNeoBench, a new benchmark designed to evaluate how well large language models understand Korean neologisms. The dataset comprises 1,785 neologisms found in online news since 2020, each accompanied by usage examples, word-formation analyses, and definitions. Experiments conducted using KoNeoBench revealed that current LLMs struggle with tasks such as identifying word components, differentiating semantic categories, and generating accurate definitions for these new words. AI
IMPACT This benchmark highlights current LLM limitations in understanding evolving language, particularly for typologically distinct languages like Korean.
RANK_REASON The cluster contains an academic paper introducing a new benchmark dataset for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →