A user analyzed their Claude Code usage and found that subagents were consuming a significant portion of their token allowance, accounting for 48.1% of all tokens and 64% of output. The high cost is attributed to the initial token load for subagent instructions and tool lists, as well as inheriting the main session's model. The user suggests two habits to mitigate this: skipping subagents for tasks that can be done in-place and pinning smaller models for subagent runs. Additionally, they highlight the impact of cache TTL, noting that extending the subagent prompt cache from five minutes to one hour can save tokens if subagents frequently idle, though it may increase costs for short bursts. AI
IMPACT Provides insights into optimizing LLM usage and managing costs for users employing agentic workflows.
RANK_REASON User analysis of a product's usage and cost implications.
Read on dev.to — Claude Code tag →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →