A 28.9 million parameter language model has been successfully run on an $8 ESP32-S3 microcontroller, achieving approximately 9 tokens per second. This significant advancement is made possible by adapting Google's Per-Layer Embeddings concept from Gemma models, allowing most of the model's parameters to reside in slow flash memory rather than fast RAM. The model, trained on the TinyStories dataset, is capable of generating short, coherent stories but lacks instruction-following or factual recall capabilities due to the limitations of its reasoning core. AI
IMPACT Demonstrates feasibility of running larger LLMs on low-power edge devices, potentially enabling new on-device AI applications.
RANK_REASON Novel application of existing LLM techniques to extremely constrained hardware.
Read on Hacker News — AI stories ≥50 points →
- esp32-ai
- slvDev
- Andrej Karpathy
- ESP32-S3
- Gemma
- llama2.c
- Microsoft Research
- Ronen Eldan
- TinyStories
- Yuanzhi Li
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →