Researchers have introduced AVShift, a new benchmark designed to evaluate authorship verification (AV) models under various distribution shifts, including genre, time, and AI-assisted writing. This benchmark, the first of its kind for German texts, comprises over 150,000 text pairs across three genres and 21 years. Experiments with AVShift indicate that fine-tuned large language models (LLMs) perform best across genres and benefit from diverse training data, while temporal drift significantly impacts AV performance. Notably, the study found no measurable AI-era distribution shift within the AVShift dataset. AI
IMPACT This benchmark will help researchers develop more robust AI writing detection systems that can account for real-world variations in text.
RANK_REASON The cluster contains an academic paper introducing a new benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →