tiktoken
PulseAugur coverage of tiktoken — every cluster mentioning tiktoken across labs, papers, and developer communities, ranked by signal.
7 day(s) with sentiment data
-
Simon Willison releases ttok 1.0, defaults to GPT-5/GPT-6 tokenizer
Simon Willison has released version 1.0 of his command-line tool, ttok, which counts tokens using OpenAI's tiktoken library. The update includes a fix for a Click warning, updated continuous integration, and a new comma…
-
tiktoken undercounts Claude tokens, causing prompt failures · 2 sources tracked
A developer discovered that OpenAI's tiktoken library, commonly used for counting tokens in LLM prompts, significantly undercounts tokens for Anthropic's Claude models. This discrepancy, which averaged 14% on the develo…
-
MCP streamlines AI agent configuration with a single source of truth
The author describes a new configuration management approach called MCP, designed to address the issue of drifting configurations across multiple AI agent clients. By centralizing server lists and skill pointers into a …
-
Developers can now count GPT tokens before API calls
Developers can now accurately count GPT tokens before sending requests to the API, preventing unexpected costs and truncation. This is achieved using the `tiktoken` library in Python and `js-tiktoken` in JavaScript, whi…
-
Developers can estimate LLM inference costs using tokenization
Developers can estimate their large language model (LLM) inference costs by modeling token usage before deploying applications. The primary challenge lies in accurately predicting input tokens, which include system prom…
-
AI Agents' Token Bills: Tool Schemas and Skills Consume Millions
A developer has identified two significant sources of token consumption in AI agent contexts: tool schemas and agent skills. The first bill, related to tool schemas, can consume over 71,000 tokens for a catalog of 255 t…
-
Developer details 6 production bugs in OpenAI-compatible streaming proxies
A developer shared insights into six specific bugs encountered while building an OpenAI-compatible proxy for streaming LLM requests. These issues, which often pass unit tests, only manifest in production environments wi…
-
Developer finds OpenAI's tiktoken library miscounts Claude tokens by up to 38%
A developer discovered that OpenAI's `tiktoken` library significantly underestimates token counts for Anthropic's Claude models, leading to unexpected API errors and budget overruns. Across 4,200 requests, `tiktoken` es…
-
Open-source Rust engine Continuum slashes AI agent token use by 96%
An open-source Rust-based memory engine named Continuum has been developed to address token limitations and memory issues in AI agents. This engine aims to reduce token costs by up to 96% and ensure complete recall acro…
-
New chunking methods improve LLM document translation quality
Researchers have developed new methods for document-level machine translation (DocMT) to overcome the limitations of current LLMs, even those with large context windows. One approach, Fixed-Range Chunking (FRC), uses dy…
-
LLM API costs soar due to quadratic token billing; sliding window offers fix
A developer highlights a common pitfall in LLM API usage where the cost of conversations can escalate quadratically due to stateless APIs requiring the resending of entire chat histories. This leads to unexpectedly high…
-
LLM tokens are an architectural constraint, not infinite resource
In production AI systems, particularly in HealthTech, LLM tokens should be treated as a finite architectural constraint rather than an infinite resource. Developers can manage this by implementing an "Estimate, Reserve,…
-
New tool visualizes LLM prompt token usage and detects waste
A new open-source tool called prompt-flamegraph has been released to help developers visualize and analyze token usage within their LLM prompts. This tool generates an interactive HTML flamegraph, breaking down token co…
-
AI agent's token usage slashed by 49% with schema optimization
A developer discovered that their AI agent, Claude Code, was consuming a significant number of tokens (8,248) before any user input, due to the large JSON schema definitions for its 48 tools being sent with every reques…
-
LLM token counting explained: why it matters for cost, context, and output
Understanding token counts is crucial for interacting with large language models, as models process text in tokens rather than words or characters. Different models and text types tokenize differently, with code and non…
-
New tool mcptoon slashes AI agent manifest costs by 70,000 tokens
A new open-source Python tool called mcptoon aims to reduce the cost and improve the performance of AI agents by optimizing how tool manifests are handled. The tool addresses the issue where clients repeatedly bill for …
-
Developer creates context compressor to prevent LLM chat memory loss
A developer has created a "Token-Aware Context Compressor" to address the issue of large language models forgetting information in long conversations. This method compresses older parts of a chat into a single summary m…
-
JSON output costs 2.6x more than CSV for LLM data, study finds
Using JSON for data output from large language models can be significantly more expensive than using CSV due to token costs associated with formatting. Pretty-printed JSON, for instance, can cost 2.6 times more than CSV…
-
New CLI tool slashes AI agent token costs by simplifying configurations
A new CLI tool called mcptoon has been developed to address inefficiencies in AI agent configurations, particularly for tools like Claude Code and Cursor. It consolidates multiple agent configuration files into a single…
-
AI agents suffer "context rot" as memory degrades with accumulated tool outputs
A developer has identified a common issue in AI agents where they forget previously provided information, a phenomenon termed "context rot." This occurs as tool outputs accumulate in the agent's context window, diluting…