PulseAugur
EN
LIVE 09:27:46

New Rushes dataset reveals GPT-5 struggles with personalized AI narrative engagement

Researchers have introduced Rushes, a new dataset and benchmark designed to study how humans engage with AI-generated interactive narratives. The dataset captures over 44,000 decision events from 8,000 users, focusing on sequential, personalized engagement rather than static judgments. Initial findings reveal a significant 'Engagement Gap,' where current large language models, including GPT-5, struggle to predict user choices, performing worse than simple baselines and classical matrix factorization techniques. This suggests that current reinforcement learning from human feedback (RLHF) methods, which often rely on population-level objectives, are insufficient for capturing the heterogeneous and context-dependent nature of user preferences in generative systems. AI

IMPACT Highlights limitations in current RLHF methods for capturing heterogeneous user preferences in interactive AI systems.

RANK_REASON The cluster contains an academic paper detailing a new dataset and benchmark for AI research. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New Rushes dataset reveals GPT-5 struggles with personalized AI narrative engagement

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Michael Xu, Jorge Leandro, Sudha Rao, Weijia Xu, Nebojsa Jojic, Gabriel DesGarennes, Chris Quirk, Bill Dolan ·

    Rushes: A Human Preference Dataset for Pluralistic Alignment

    arXiv:2607.20767v1 Announce Type: new Abstract: We introduce Rushes, a dataset and benchmark for studying revealed human engagement preferences in interactive narrative environments. Rushes is collected through a game interface where users interact with AI-generated branching nar…