A developer discovered that OpenAI's tiktoken library, commonly used for counting tokens in LLM prompts, significantly undercounts tokens for Anthropic's Claude models. This discrepancy, which averaged 14% on the developer's code-heavy inputs, led to numerous "prompt too long" errors when requests neared Claude's context limit. The issue arises because tiktoken uses OpenAI's tokenization vocabulary, which differs from Claude's. The developer found that specific content types like JSON and Python code exacerbated the undercounting. A proposed solution involves an initial estimate with tiktoken and a multiplier, followed by a precise count using Anthropic's `count_tokens` endpoint to trim prompts as needed. AI
IMPACT Highlights the need for precise token counting for LLM API calls, impacting cost estimation and prompt engineering.
RANK_REASON Developer discovers a practical issue with a common tool (tiktoken) when used with a specific AI model (Claude), leading to a workaround.
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →