Researchers have evaluated the effectiveness of various vision-language models (VLMs) in grading handwritten examinations for outcome-based education. The study compared configurations including Qwen2.5-VL, InternVL3, and Pixtral, utilizing methods like zero-shot prompting, few-shot prompting, and Low-Rank Adaptation (LoRA). Qwen2.5-VL with LoRA achieved a Quadratic Weighted Kappa (QWK) of 0.727, surpassing the average human-pair QWK of 0.551, indicating potential for automated grading. However, the research also highlighted challenges such as mark variability across runs and a lack of consensus on the usefulness of model-generated explanations. AI
IMPACT These models show potential for automating the grading of handwritten exams, improving efficiency and consistency over manual methods.
RANK_REASON The cluster is based on an academic paper detailing research into the application of vision-language models for a specific task. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- CatalyzeX
- DagsHub
- Donut
- Gotit.pub
- Hugging Face
- InternVL3
- LoRA+
- Pixtral
- quadratic weighted kappa
- Qwen2.5-VL
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →