PulseAugur
EN
LIVE 07:08:53

New benchmark tackles long-form LLM text attribution across languages

Researchers have introduced MultiGhostBench, a new multilingual benchmark designed to evaluate the attribution of long-form text generated by Large Language Models (LLMs). This benchmark includes 928 books across six languages and three scripts, with an average length of approximately 59,000 words, and supports evaluation under various distribution shifts like domain, author, and language changes. Initial evaluations indicate that current attribution methods struggle to perform consistently across different settings, with performance degrading under distribution shifts, though Transformer-based detectors show some cross-lingual capabilities. AI

IMPACT This benchmark could drive advancements in detecting AI-generated content, crucial for academic integrity and combating misinformation.

RANK_REASON The cluster describes a new academic paper introducing a benchmark for LLM text attribution. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.IR (Information Retrieval) →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

New benchmark tackles long-form LLM text attribution across languages

How we ranked this

Signal score
2 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
The cluster describes a new academic paper introducing a benchmark for LLM text attribution. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [2]

  1. arXiv cs.AI TIER_1 English(EN) · Matteo Greco, Anudeex Shetty, Andrea Tagarelli, Jey Han Lau ·

    MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts

    arXiv:2609.02379v1 Announce Type: cross Abstract: While existing work on LLM authorship attribution (AA) has made progress, available benchmarks remain limited, often focusing on English, controlled settings, or relatively outdated models, with the few multilingual studies consid…

  2. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Jey Han Lau ·

    MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts

    While existing work on LLM authorship attribution (AA) has made progress, available benchmarks remain limited, often focusing on English, controlled settings, or relatively outdated models, with the few multilingual studies considering only relatively short texts. We introduce Mu…