Researchers have introduced MultiGhostBench, a new multilingual benchmark designed to evaluate the attribution of long-form text generated by Large Language Models (LLMs). This benchmark includes 928 books generated by five recent LLMs across six languages and three scripts, with an average length of approximately 59,000 words per book. Evaluations indicate that current attribution methods struggle, particularly under distribution shifts, and that while Transformer-based detectors show some cross-lingual capabilities, statistical and fingerprint-based detectors are more language-dependent. AI
IMPACT This benchmark will aid in developing more robust methods for identifying AI-generated text, crucial for combating misinformation and ensuring content authenticity.
RANK_REASON The item describes a new academic paper introducing a benchmark for LLM-generated text attribution. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →