PulseAugur
EN
LIVE 18:19:25

LLM efficiency measured by tokens per Big Mac calorie

SemiAnalysis has introduced a novel metric for evaluating large language model (LLM) inference efficiency: tokens per second per megawatt (tok/s/MW). This metric allows for a comparison of energy consumption against output, with calculations showing that a single Big Mac's worth of calories could theoretically produce approximately 18,000 tokens on a B300 GPU configuration. In contrast, the same caloric energy could power human speech to produce around 400,000 tokens, highlighting a significant difference in energy efficiency between current LLM technology and human biological processes. AI

IMPACT Introduces a new metric for evaluating LLM energy efficiency, highlighting the gap between current AI capabilities and human biological energy efficiency.

RANK_REASON The cluster discusses a novel metric for LLM efficiency and compares it to human speech, but does not announce a new model, research breakthrough, or significant industry event.

Read on X — SemiAnalysis →

AI-generated summary · Google Gemini · from 4 sources. How we write summaries →

LLM efficiency measured by tokens per Big Mac calorie

COVERAGE [4]

  1. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    Unfortunately, the average human is much dumber than most frontier language models. So while we can speak to throughput per Big Mac, goodput might be a differen

    Unfortunately, the average human is much dumber than most frontier language models. So while we can speak to throughput per Big Mac, goodput might be a different story 😅 (4/4) https://t.co/FV2YZUeCPT

  2. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    Finally, a Big Mac contains roughly 580 calories which means:

    Finally, a Big Mac contains roughly 580 calories which means: 🟠 B300: 30.8 × 580 = ~18,000 output tokens per Big Mac 🟠 Human: 690 × 580 = ~400,000 output tokens per Big Mac One Big Mac can produce a lot of output tokens! (3/4) https://t.co/0LY7kD9rDT

  3. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    Using the fact that 1 MW = 1,000,000 J/s and 1 joule = 0.000239 food Calories, we can convert power efficiency into tokens per Calorie:

    Using the fact that 1 MW = 1,000,000 J/s and 1 joule = 0.000239 food Calories, we can convert power efficiency into tokens per Calorie: 7,368 tok/s/MW = 0.007368 tok/J 0.007368 ÷ 0.000239 = 30.8 output tok/Calorie Given that an average human speaks at ~3.3 tok/s while the http…

  4. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    A useful metric for evaluating LLM inference is tok/s/MW: how many tokens a system generates per second for each megawatt of all-in provisioned power.

    A useful metric for evaluating LLM inference is tok/s/MW: how many tokens a system generates per second for each megawatt of all-in provisioned power. At one concurrent user, the B300 configuration below generates ~14 output tok/s per GPU on DeepSeek V4. The SemiAnalysis https:/…