PulseAugur
EN
LIVE 10:54:17

New LLM tools evaluate essay scoring bias and student AI reliance

Researchers have developed new tools and analyses to evaluate the performance and fairness of large language models (LLMs) in academic writing. One study introduces WrAFT, a modular system for automated essay scoring and feedback that achieves state-of-the-art performance using models like Llama 3.3 70B Instruct and GPT-4o. Another paper investigates first-language bias in an LLM adapted for automated essay scoring on TOEFL essays, finding that essays from European-language backgrounds received higher scores than those from East-Asian backgrounds, despite stable cross-prompt generalization. Additionally, a new scale, GenAI-RTS, has been developed and validated to measure how students rely on generative AI in academic writing, categorizing reliance into strategic, instrumental, dependent, and dialogic types. AI

IMPACT These studies offer new tools and insights for evaluating LLM fairness and understanding student interaction with AI in academic writing.

RANK_REASON Multiple academic papers published on arXiv detailing new research into LLM applications for automated essay scoring, bias detection, and student AI reliance.

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 8 sources. How we write summaries →

New LLM tools evaluate essay scoring bias and student AI reliance

COVERAGE [8]

  1. arXiv cs.CL TIER_1 English(EN) · John Maurice Gayed ·

    Investigating first-language bias in LLM-based automated essay scoring: A cross-prompt evaluation of an open-weight AI-model on TOEFL essays

    arXiv:2607.14605v1 Announce Type: new Abstract: This study examines the cross-prompt generalization and first-language (L1) scoring effects of a LoRA-adapted open-weight large language model (Gemma-3-27B-it) applied to automated essay scoring. Using the identical model and infere…

  2. arXiv cs.AI TIER_1 English(EN) · Shahin Hossain, Tukhbita Afroz Nawmi ·

    Measuring How Students Rely on Generative AI in Academic Writing: Development and Multi-Source Validation of the Generative AI Reliance Types Scale (GenAI-RTS)

    arXiv:2607.14301v1 Announce Type: new Abstract: As generative AI (GenAI) becomes increasingly embedded in undergraduate academic writing, how students rely on these tools, rather than simply whether they use them, has become a central question for learning, academic integrity, an…

  3. arXiv cs.AI TIER_1 English(EN) · Adnan Labib, Yixuan Huang, Jiahui Wu, John Maurice Gayed, Zheng Yuan, Qiao Wang ·

    WrAFT: a Modularized Automated Writing Evaluation System for Argumentative Essays

    arXiv:2607.14524v1 Announce Type: new Abstract: This study presents WrAFT, a Writing Assessment and Feedback Tool, that delivers both accurate and reliable scores and effective comprehensive feedback to argumentative essays. WrAFT adopts a modular design by dividing automated wri…

  4. arXiv cs.CL TIER_1 English(EN) · Steven Coyne, Diana Galvan-Sosa, Ryan Spring, Machi Shimmei, Michael Zock, Keisuke Sakaguchi, Kentaro Inui ·

    How Well Does AI-Generated Feedback Work? Intrinsic and Extrinsic Evaluation across more than 20,000 EFL Essay Drafts

    arXiv:2607.14591v1 Announce Type: new Abstract: This study examines feedback in English as a Foreign Language (EFL) writing contexts, focusing on written corrective feedback (WCF). Large language models (LLMs) can provide WCF at scale, but aligning them with pedagogical best prac…

  5. arXiv cs.CL TIER_1 English(EN) · John Maurice Gayed ·

    Investigating first-language bias in LLM-based automated essay scoring: A cross-prompt evaluation of an open-weight AI-model on TOEFL essays

    This study examines the cross-prompt generalization and first-language (L1) scoring effects of a LoRA-adapted open-weight large language model (Gemma-3-27B-it) applied to automated essay scoring. Using the identical model and inference configuration reported in "AiAWE: An Open-So…

  6. arXiv cs.CL TIER_1 English(EN) · Kentaro Inui ·

    How Well Does AI-Generated Feedback Work? Intrinsic and Extrinsic Evaluation across more than 20,000 EFL Essay Drafts

    This study examines feedback in English as a Foreign Language (EFL) writing contexts, focusing on written corrective feedback (WCF). Large language models (LLMs) can provide WCF at scale, but aligning them with pedagogical best practices remains an ongoing challenge. WCF meeting …

  7. arXiv cs.CL TIER_1 English(EN) · Qiao Wang ·

    WrAFT: a Modularized Automated Writing Evaluation System for Argumentative Essays

    This study presents WrAFT, a Writing Assessment and Feedback Tool, that delivers both accurate and reliable scores and effective comprehensive feedback to argumentative essays. WrAFT adopts a modular design by dividing automated writing evaluation (AWE) tasks into scoring, surfac…

  8. arXiv cs.CL TIER_1 English(EN) · Tukhbita Afroz Nawmi ·

    Measuring How Students Rely on Generative AI in Academic Writing: Development and Multi-Source Validation of the Generative AI Reliance Types Scale (GenAI-RTS)

    As generative AI (GenAI) becomes increasingly embedded in undergraduate academic writing, how students rely on these tools, rather than simply whether they use them, has become a central question for learning, academic integrity, and educational equity. Existing measures of relia…