PulseAugur
EN
LIVE 14:14:20

New Llama-based model improves essay scoring generalization to unseen rubrics

Researchers have developed a new framework for automated essay scoring (AES) that improves generalization to unseen scoring rubrics. By using rubric-agnostic intermediate representations called 'traits' and controlled supervision, their fine-tuned Llama-based model achieved a 5.0% improvement in macro F1 score compared to a baseline without traits in the most challenging scenario. This approach also showed competitive performance against proprietary models, with the best open-source model outperforming GPT-5-mini by 2.1% macro F1 and trailing GPT-5 by only 1.9%. AI

IMPACT Enhances the ability of AI systems to evaluate essays across different grading criteria, potentially improving educational tools.

RANK_REASON Academic paper detailing a new method for automated essay scoring. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New Llama-based model improves essay scoring generalization to unseen rubrics

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Andrew Lan ·

    When Rubrics Change: Cross-Rubric Generalization for Critical Thinking Essay Scoring

    Automated essay scoring (AES) research has largely focused on cross-prompt generalization, where essays from unseen prompts are scored while the scoring criteria are typically held constant. In practice, however, educators may revise or even introduce new rubrics in their scoring…