A new study evaluated the inter-rater reliability of large language models (LLMs) and rule-based systems in annotating inferential narrative features within a Turkish corpus. The research found that models like Gemini 2.5 Flash, Grok, Claude Fable-5, and ChatGPT 5.5, along with a rule-based detector, showed low agreement with human annotators on features such as materialized metaphor. The study suggests that these inferential features may be too complex for current automatic detection or that the definitions themselves are not yet operational enough for consistent application by any rater. AI
IMPACT Highlights limitations of current LLMs in nuanced text analysis, suggesting a gap in understanding complex inferential features.
RANK_REASON Academic paper evaluating LLM performance on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →