A user discovered that their reported usage of 1.36 billion tokens in Claude Code over a week was primarily due to prompt caching, with only about 36 million tokens representing new input or generated output. This high cache read ratio, accounting for 97.4% of the total, indicates efficient reuse of context for an agent-based system like Claude Code. While cost-effective, the user notes that high cache efficiency does not directly equate to improved inference quality or reasoning accuracy. AI
IMPACT Highlights the importance of understanding LLM token usage metrics, particularly for agentic systems, and the cost implications of prompt caching.
RANK_REASON User-provided analysis of a product's usage metrics, not a direct product release or research finding.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →