PulseAugur
中
实时 01:16:12
English(EN) Counting Claude tokens with tiktoken undercounted my prompts by 14%. 37 of 418 requests failed. Here's the audit and the count_tokens fix. # claude # llm # pyth

tiktoken 低估 Claude 的 token 数量,导致提示词失败 · 跟踪 2 个来源

一位开发者发现,OpenAI 的 tiktoken 库(常用于计算 LLM 提示词的 token 数量)在计算 Anthropic 的 Claude 模型时,token 数量被严重低估。这种差异在开发者代码量大的输入上平均为 14%,当请求接近 Claude 的上下文限制时,导致了许多“提示词过长”的错误。问题在于 tiktoken 使用的是 OpenAI 的分词词汇表,这与 Claude 的不同。开发者发现 JSON 和 Python 代码等特定内容类型加剧了低估。提出的解决方案是先使用 tiktoken 进行估算并乘以一个系数,然后使用 Anthropic 的 `count_tokens` 端点进行精确计数,以根据需要截断提示词。 AI

影响 强调了对 LLM API 调用进行精确 token 计数的必要性,影响成本估算和提示词工程。

排序理由 开发者在使用特定 AI 模型(Claude)时发现了一个常见工具(tiktoken)的实际问题,并找到了一个解决方法。

在 Mastodon — mastodon.social 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

tiktoken 低估 Claude 的 token 数量,导致提示词失败 · 跟踪 2 个来源

本文如何被排名

Signal score
4 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
开发者在使用特定 AI 模型(Claude)时发现了一个常见工具(tiktoken)的实际问题,并找到了一个解决方法。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
product, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准。

报道来源 [2]

  1. dev.to — LLM tag TIER_1 English(EN) · jidonglab ·

    tiktoken 低估 Claude Token 数量:418 次请求中有 37 次超出限制

    <p>My nightly digest job died at 2:14 a.m. with a 400 and the message <code>prompt is too long</code>. My packer had measured the prompt at 181,874 tokens, comfortably under the 200K window. Claude said it was 207,431.</p> <p>That's a 14% gap, and it came from one lazy decision I…

  2. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    使用 tiktoken 计算 Claude token 低估了我的提示 14%。418 次请求中有 37 次失败。这是审计和 count_tokens 的修复方法。# claude # llm # pyth

    Counting Claude tokens with tiktoken undercounted my prompts by 14%. 37 of 418 requests failed. Here's the audit and the count_tokens fix. # claude # llm # python # ai # software # coding # development # engineering # inclusive # community tiktoken Undercounts Claude Tokens: 37 o…