A new study published on arXiv explores how large language models (LLMs) used in student assessment are influenced by demographic signals. Researchers found that LLMs can alter their scoring, feedback, and answers based on both explicit mentions of demographics and implicit cues from conversation history. While LLMs adjusted readability for explicit educational levels, implicit conditions led to unpredictable biases, such as lower sentiment scores for responses from lower-education backgrounds in question answering tasks. The findings highlight the demographic sensitivity of LLMs in educational assessment. AI
IMPACT Highlights potential biases in AI-driven educational tools, necessitating careful development and deployment to ensure fairness.
RANK_REASON Academic paper detailing research findings on LLM behavior. [lever_c_demoted from research: ic=1 ai=1.0]
- Automated Essay Scoring
- formative assessment
- large-language models
- Metalinguistic Question Answering
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →