A new benchmark called LiveMacroEval has been developed to assess the real-time macroeconomic nowcasting capabilities of large language models (LLMs). This benchmark is designed to be contamination-resistant, evaluating LLM agents' hourly nowcasts for sixteen U.S. macroeconomic indicators during their pre-release window. The LLMs' performance is compared against institutional forecasts, professional consensus, and an ARIMA baseline, with aggregate accuracy found to be comparable across these benchmarks. AI
IMPACT LLMs demonstrate potential for real-time economic forecasting, comparable to expert human analysis.
RANK_REASON The cluster describes a new academic paper introducing a novel benchmark for evaluating LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
- Auto ARIMA
- Bloomberg ECOS
- consumer price index
- Federal Reserve System
- gross domestic product
- Hugging Face
- Large language model
- LiveMacroEval
- Polymarket
- U.S.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →