PulseAugur
EN
LIVE 01:19:35

28.9M parameter LLM runs on $8 microcontroller using Gemma-inspired technique

A 28.9 million parameter language model has been successfully run on an $8 ESP32-S3 microcontroller, achieving approximately 9 tokens per second. This significant advancement is made possible by adapting Google's Per-Layer Embeddings concept from Gemma models, allowing most of the model's parameters to reside in slow flash memory rather than fast RAM. The model, trained on the TinyStories dataset, is capable of generating short, coherent stories but lacks instruction-following or factual recall capabilities due to the limitations of its reasoning core. AI

IMPACT Demonstrates feasibility of running larger LLMs on low-power edge devices, potentially enabling new on-device AI applications.

RANK_REASON Novel application of existing LLM techniques to extremely constrained hardware.

Read on Hacker News — AI stories ≥50 points →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

28.9M parameter LLM runs on $8 microcontroller using Gemma-inspired technique

COVERAGE [2]

  1. Hacker News — AI stories ≥50 points TIER_1 English(EN) · boveyking ·

    Running a 28.9M parameter LLM on an $8 microcontroller

  2. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Running a 28.9M parameter LLM on an $8 microcontroller https:// github.com/slvDev/esp32-ai # ai # github # llm

    Running a 28.9M parameter LLM on an $8 microcontroller https:// github.com/slvDev/esp32-ai # ai # github # llm