A recent analysis of the Jev language model revealed that it does not possess memory of past events, contrary to initial expectations. Despite being trained on vast amounts of text, Jev failed to recall outcomes from prediction markets resolved up to 20 months before its release, performing only slightly better than a coin flip on these historical data points. While Jev showed a marginally higher score on a market related to Donald Trump's inauguration, its performance across other historical markets, including those concerning Federal Reserve actions, remained consistently low, suggesting a lack of true recall. AI
IMPACT This finding suggests that current LLMs may not effectively retain knowledge of past events, impacting their reliability for historical prediction or recall tasks.
RANK_REASON Analysis of an LLM's performance on historical data. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →