PulseAugur
EN
LIVE 08:23:51

Study finds AI thesis assessment struggles to match human priorities

A study surveyed 84 thesis supervisors across four academic disciplines to understand their prioritization of thesis assessment criteria. The research found significant differences between the supervisors' derived criterion weights and the default weights used by the AI assessment system RubiSCoT. When these supervisor-derived weights were integrated into RubiSCoT, the AI-generated evaluations showed only a marginal improvement in alignment with human assessments, indicating that criterion-weight calibration alone is insufficient to bridge the gap. AI

IMPACT This research highlights the challenges in aligning AI assessment tools with nuanced human evaluation criteria, suggesting further work is needed beyond simple weight calibration.

RANK_REASON The cluster contains an academic paper detailing an empirical study on AI-based thesis assessment. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Study finds AI thesis assessment struggles to match human priorities

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Garv Vikram Gursahaney, Baskhad Idrisov, Thorsten Fr\"ohlich, Tim Schlippe ·

    AI-Based Thesis Assessment: An Empirical Study of Human Evaluation Priorities and Their Impact on Automated Assessment

    arXiv:2608.00717v1 Announce Type: cross Abstract: Rubric-based AI systems for thesis assessment use criterion weights to assign different levels of importance to evaluation criteria. These weights are typically defined through expert judgment, although little empirical evidence e…