PulseAugur
EN
LIVE 06:34:10

New benchmark reveals AI-text detectors struggle with rewritten human content

A new benchmark dataset called ARB has been developed to evaluate the effectiveness of AI-text detectors when human-authored content is rewritten by large language models. The dataset includes human-written text, direct LLM generation, and LLM-rewritten versions of both human and LLM text, utilizing four open-weight generators. Evaluations showed that detectors performed significantly worse on LLM-rewritten human text compared to direct LLM generation, indicating a gap in current detection capabilities. AI

IMPACT Highlights a critical vulnerability in current AI-text detection methods when faced with paraphrased or rewritten content.

RANK_REASON The item is a research paper introducing a new benchmark dataset for evaluating AI-text detectors. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New benchmark reveals AI-text detectors struggle with rewritten human content

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Gaetano Perrone, Simon Pietro Romano ·

    ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation

    arXiv:2607.29539v1 Announce Type: cross Abstract: Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs). While prior work has shown that rewriting and paraphrasing can degrade detector performance, it rema…