Researchers have developed HouseholdBench, a new evaluation dataset designed to assess how well large language models (LLMs) can predict household economic behavior. The benchmark combines data from six U.S. household surveys, covering 32 prediction tasks related to consumption, income, labor, expectations, and housing. Initial evaluations show that most LLMs outperform a basic baseline, with the best models reducing error by over 12% on numeric outcomes, though gradient-boosted tree models generally perform better. The study also found that fine-tuning and aggregating predictions from a 4-billion parameter open-weight model can significantly improve its performance to match that of proprietary LLMs. AI
IMPACT Establishes a new standard for evaluating LLM capabilities in economic prediction, potentially guiding future model development for socio-economic applications.
RANK_REASON Academic paper introducing a new benchmark dataset for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →