A study by Heatseeker CEO Kate O'Keeffe tested large language models ChatGPT, Claude, and Gemini against 11 past advertising experiments to see if they could predict customer response. The models collectively performed poorly, with all three being incorrect in six of the tests and selecting the same incorrect answer in four instances. This research suggests that while LLMs can generate plausible explanations, their ability to accurately predict observed customer behavior based on provided data is limited, raising questions about the trustworthiness of their outputs in real-world applications. AI
IMPACT Highlights the gap between LLM plausibility and real-world predictive accuracy for customer behavior.
RANK_REASON Opinion piece by a CEO discussing limitations of LLMs based on a small study.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →