A new benchmark dataset called ARB has been developed to evaluate the effectiveness of AI-text detectors when human-authored content is rewritten by large language models. The dataset includes human-written text, direct LLM generation, and LLM-rewritten versions of both human and LLM text, utilizing four open-weight generators. Evaluations showed that detectors performed significantly worse on LLM-rewritten human text compared to direct LLM generation, indicating a gap in current detection capabilities. AI
IMPACT Highlights a critical vulnerability in current AI-text detection methods when faced with paraphrased or rewritten content.
RANK_REASON The item is a research paper introducing a new benchmark dataset for evaluating AI-text detectors. [lever_c_demoted from research: ic=1 ai=1.0]
- AI-text detector
- BERT-Defense
- Binoculars-falcon-7b
- FastDetectGPT
- Gaetano Perrone Mr.
- Gemma-2-9B
- large language models
- Llama-3.2-3B
- Mistral-7B
- OpenWebText
- Qwen2.5-7B
- RADAR
- RoBERTa-Defense
- WritingPrompts
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →