Large language models from OpenAI and Anthropic are becoming more capable, but users are finding that their bills are increasing despite lower per-token prices. This is because the advanced models perform more internal "reasoning" or "thinking" steps, which consume a significant number of tokens that are billed as output, even though they are not visible to the user. For example, a model might provide a 200-token answer, but the underlying reasoning process could have used 12,000 tokens. This increased cost is often justified by the models' improved ability to understand and execute complex or ambiguous prompts, but users are advised to monitor the `reasoning_tokens` or `thinking_tokens` fields to understand where their costs are originating. AI
IMPACT Users should monitor internal reasoning token usage to manage costs, as advanced LLMs perform more complex, billed computations.
RANK_REASON User-written analysis of pricing and usage patterns for LLM APIs.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →