Researchers have developed new tools and analyses to evaluate the performance and fairness of large language models (LLMs) in academic writing. One study introduces WrAFT, a modular system for automated essay scoring and feedback that achieves state-of-the-art performance using models like Llama 3.3 70B Instruct and GPT-4o. Another paper investigates first-language bias in an LLM adapted for automated essay scoring on TOEFL essays, finding that essays from European-language backgrounds received higher scores than those from East-Asian backgrounds, despite stable cross-prompt generalization. Additionally, a new scale, GenAI-RTS, has been developed and validated to measure how students rely on generative AI in academic writing, categorizing reliance into strategic, instrumental, dependent, and dialogic types. AI
IMPACT These studies offer new tools and insights for evaluating LLM fairness and understanding student interaction with AI in academic writing.
RANK_REASON Multiple academic papers published on arXiv detailing new research into LLM applications for automated essay scoring, bias detection, and student AI reliance.
- AiAWE
- arXiv
- Gemma 3 27B IT
- GenAI-RTS
- GPT-4o
- Hugging Face
- John Maurice Gayed
- large language models
- Llama 3.3 70B Instruct
- TOEFL
AI-generated summary · Google Gemini · from 8 sources. How we write summaries →