A new paper explores how different reward functions, known as proper scoring rules, impact the forecasting abilities and behaviors of large language models. While these rules theoretically encourage honest probability reporting, the study found that models trained with different rules exhibited variations in calibration, probability usage, and estimated profiles of bias, information, and noise. The Brier-trained model achieved the best Brier score and AUC-ROC, whereas the log-trained model excelled in log score and calibration error, indicating that the choice of reward function significantly shapes not only forecasting accuracy but also the structure of forecasting errors. AI
IMPACT This research highlights that the choice of training objective for LLMs can significantly alter their forecasting capabilities and error structures, suggesting a need for careful consideration of reward functions in developing reliable AI forecasters.
RANK_REASON The item is an academic paper detailing research into LLM behavior. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →