A recent analysis of large language models revealed a significant disparity in task costs, with a 10.6x spread observed across GPT, Claude, Gemini, and Kimi. This wide cost variation occurred despite the models' published rates differing by only 2x. The study found that invisible reasoning tokens, which are billed at the output rate but not displayed in the response, were a primary driver of these cost differences. For instance, one model used 197 invisible tokens for a simple one-word classification task. AI
IMPACT Highlights the hidden costs of LLM usage, emphasizing the need for better cost optimization strategies beyond published rates.
RANK_REASON Analysis of LLM performance and cost. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →