A new research paper published on arXiv examines the reliability and consistency of large language models (LLMs) when used as essay graders. The study, which treated LLMs as human raters, found significant variations in severity and version instability across different LLM judges and versions. While LLMs showed some self-consistency, their accuracy did not reach human levels, and their tendency to exhibit 'halo' effects was comparable to trained human raters when calibrated appropriately. AI
IMPACT Highlights potential issues with LLM reliability in educational assessment, suggesting caution in their deployment for grading.
RANK_REASON Research paper published on arXiv detailing findings about LLM performance. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →