A developer has successfully run a 28.9 million parameter language model on an $8 ESP32-S3 microcontroller, achieving approximately 9 tokens per second without cloud dependency. This significant advancement in edge AI leverages Google's Per-Layer Embeddings technique, allowing most of the model's parameters to reside in slow flash memory while keeping the core processing components in the chip's limited SRAM. The model, trained on the TinyStories dataset, generates short, coherent narratives, demonstrating a new capability for low-cost, offline generative AI applications. AI
IMPACT Enables sophisticated AI capabilities on extremely low-cost, power-efficient edge devices, opening new possibilities for offline smart sensors and embedded systems.
RANK_REASON Project demonstrates running a large LLM on commodity hardware, which is a significant tooling advancement for edge AI.
Read on Mastodon — fosstodon.org →
- esp32-ai
- slvDev
- Andrej Karpathy
- ESP32-S3
- Gemma
- llama2.c
- Microsoft Research
- Ronen Eldan
- TinyStories
- Yuanzhi Li
- GitHub
- Hacker News
- Per-Layer Embeddings
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →