A new paper from researchers including Lukas Gehring explores the limitations of current Large Language Model (LLM) text detectors in educational settings. The study highlights that these detectors often fail to accurately distinguish between human-written text and text generated with varying degrees of AI assistance, particularly at intermediate contribution levels. To address this, the researchers propose a contribution-aware evaluation framework and introduce GEDE, a new benchmark dataset with over 12,500 generated essays, to better model realistic human-AI collaboration scenarios and assess detection systems across different policies and models. AI
IMPACT Current LLM text detectors are unsuitable for reliably enforcing academic integrity policies due to their inability to accurately classify AI-assisted writing.
RANK_REASON The cluster contains a research paper detailing a new evaluation framework and benchmark dataset for LLM text detection in education. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Generative Essay Detection in Education
- Gotit.pub
- Hugging Face
- Large Language Models
- Lukas Gehring
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →