Researchers have introduced MultiGhostBench, a new multilingual benchmark designed to evaluate the attribution of long-form text generated by Large Language Models (LLMs). This benchmark includes 928 books across six languages and three scripts, with an average length of approximately 59,000 words, and supports evaluation under various distribution shifts like domain, author, and language changes. Initial evaluations indicate that current attribution methods struggle to perform consistently across different settings, with performance degrading under distribution shifts, though Transformer-based detectors show some cross-lingual capabilities. AI
IMPACT This benchmark could drive advancements in detecting AI-generated content, crucial for academic integrity and combating misinformation.
RANK_REASON The cluster describes a new academic paper introducing a benchmark for LLM text attribution. [lever_c_demoted from research: ic=1 ai=1.0]
Read on arXiv cs.IR (Information Retrieval) →
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- LLM
- MultiGhostBench
- ScienceCast
- Transformer++
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →