PulseAugur
EN
LIVE 22:54:51

LLM test-time reasoning fails to improve financial trading returns, study finds

A new study published on arXiv investigates whether the increased computational cost of test-time reasoning in large language models (LLMs) translates into improved economic outcomes in financial trading. Researchers analyzed LLMs from the DeepSeek, GPT, and Gemini families, varying reasoning effort across numerical, news-based, and masked news inputs for U.S. equities over a full year. The findings indicate that additional reasoning does not consistently enhance net portfolio returns, and for DeepSeek models, performance was non-monotonic with increased reasoning. Repeated generations also led to unstable portfolio selections, suggesting that the economic value of enhanced reasoning needs task-specific validation before deployment. AI

IMPACT Suggests that advanced reasoning capabilities in LLMs may not directly translate to improved financial decision-making, highlighting the need for task-specific validation.

RANK_REASON Academic paper on LLM capabilities and limitations. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

LLM test-time reasoning fails to improve financial trading returns, study finds

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Jiayi Chen, Guiling Wang ·

    The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?

    arXiv:2609.30705v1 Announce Type: new Abstract: While inference-time reasoning in large language models (LLMs) promises better decision making, its higher computational cost may not yield better economic outcomes. Yet reasoning controls are rarely evaluated as economic interventi…