Researchers have introduced Rushes, a new dataset and benchmark designed to study how humans engage with AI-generated interactive narratives. The dataset captures over 44,000 decision events from 8,000 users, focusing on sequential, personalized engagement rather than static judgments. Initial findings reveal a significant 'Engagement Gap,' where current large language models, including GPT-5, struggle to predict user choices, performing worse than simple baselines and classical matrix factorization techniques. This suggests that current reinforcement learning from human feedback (RLHF) methods, which often rely on population-level objectives, are insufficient for capturing the heterogeneous and context-dependent nature of user preferences in generative systems. AI
IMPACT Highlights limitations in current RLHF methods for capturing heterogeneous user preferences in interactive AI systems.
RANK_REASON The cluster contains an academic paper detailing a new dataset and benchmark for AI research. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- GPT-5
- Hugging Face
- reinforcement learning from human feedback
- Rushes
- singular value decomposition
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →