PulseAugur
EN
LIVE 05:18:19

New benchmark reveals LLMs fail to achieve human-level authorship

A new benchmark called PersonalBench has been developed to measure the effectiveness of large language model (LLM) personalization. The benchmark evaluates LLM outputs based on three criteria: authorship verification using a trained model (LUAR), an LLM-as-judge approach, and automated stylometrics. Experiments across 50 authors and two model families, Qwen 3 and GLM-4, revealed that while personalization methods can differentiate output styles, they do not bridge the gap to human authorship. The LLM's own stylistic fingerprint remains dominant, with generated text being more distant from human authors than humans are from each other. AI

IMPACT This benchmark could drive future research in LLM personalization by providing a standardized way to measure true authorship mimicry.

RANK_REASON The cluster describes a new academic paper introducing a benchmark for evaluating LLM personalization. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New benchmark reveals LLMs fail to achieve human-level authorship

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Yash Ganpat Sawant ·

    PersonalBench: Measuring the Authorship Gap in LLM Personalization

    arXiv:2608.19746v1 Announce Type: new Abstract: Personalized text generation aims to make LLMs write in a specific individual's style, yet existing benchmarks measure task accuracy or preference alignment rather than whether the model's output actually resembles the target author…