A new benchmark called THPT-Ladder has been developed to evaluate language models on human exams, specifically addressing the issue of partial credit in grading. This benchmark uses Vietnam's 2025 Convex Marking Scheme, which deviates from standard accuracy metrics by not awarding proportional credit for partially correct answers. The study found that this grading scheme significantly penalizes models, altering their perceived competence and percentile rankings compared to human candidates. The research highlights that traditional accuracy measures do not adequately capture a model's performance under such non-additive grading systems. AI
IMPACT Highlights the need for AI evaluation methods that account for nuanced grading schemes, impacting how AI performance is understood in educational contexts.
RANK_REASON The cluster contains an academic paper introducing a new benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →