A study called Humor Arena evaluated 20 large language models on their ability to generate jokes, using a dataset of 360 joke prompts. Fable 5 emerged as the funniest, achieving an estimated score of 66.8 points out of 100, with Fable 5.1 scoring 58.2. The evaluation utilized an automated judge, which was audited against human preferences and found to correlate highly with them. AI
IMPACT This study provides insights into the humor generation capabilities of LLMs, potentially influencing future model development and fine-tuning for more engaging user interactions.
RANK_REASON The cluster describes the results of a benchmark study evaluating LLM humor generation capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →