
A 28.9M PARAMETER LANGUAGE MODEL RUNS ENTIRELY ON A CHEAP ESP32-S3 MICROCONTROLLER
A 28.9 million parameter model runs entirely on a cheap ESP32-S3 microcontroller with no server behind it.
by slvDev
AIOpen-hardware
- difficulty
- ●●●○○
- time
- a weekend
- license
- MIT
- repo
- repo ACTIVE4,458 stars
●●●○○ · a weekend · MIT · 4,458 stars · repo ACTIVE
WHAT YOU’LL NEED
Jump to section
COMPAREE VERDICT
This is a striking demonstration of a twenty-nine-million parameter language model running entirely on a microcontroller with no connectivity. The ESP32-S3 has 512KB of SRAM and 8MB of PSRAM, so the model should not fit — but slvDev borrowed Per-Layer Embeddings from Google's Gemma 3n and parked 25 million parameters in a lookup table in slow flash. Each token reads roughly 450 bytes out, activations stay in SRAM, and the dense core sits in PSRAM. The whole 28.9M parameter model quantized to 4 bits is 14.9MB. It generates at roughly 9.5 to 9.9 tokens per second end to end, displayed on a small screen. The measured benefit over a same-core baseline that fits in SRAM is 0.098 nats, about 9.3 percent perplexity improvement. For scale, earlier independent work on the same chip family ran 260 thousand parameters, so this is roughly 110 times larger. The honest limit is that the published language models were trained on TinyStories or a narrow espresso Q&A dataset — they will not answer general questions, follow instructions, write code, or know facts. If you want to understand how small models are deployed on constrained hardware or explore architectural tricks at the memory boundary, this is a very clean example. If you want a general-purpose assistant, this is not it and training your own will take more than a weekend. The single thing most likely to go wrong is expecting ChatGPT behaviour from a 28.9M parameter TinyStories model.
IN THE REPO
GOOD TO KNOW
- —MIT licence permits commercial use.
- —Repository contains firmware, training and export code, and three published models on Hugging Face (TinyStories, Barista and the fruit fly connectome demo).
- —No formal bill of materials: the README names the chip (ESP32-S3 N16R8: 8MB PSRAM, 16MB flash) and the firmware drives a small SH110X OLED display.
- —No CAD files — this is a software project with off-the-shelf hardware.
- —Documentation covers flashing, model conversion from PyTorch, and the Per-Layer Embeddings architecture.
- —The trained models are TinyStories-based: they write short simple stories, not general-purpose text.
Parts to buy
3 itemsFrom our check of the build. Exact quantities and part numbers are in the creator’s BOM.
Can I build this?
Build at your own risk. Projects involve tools, electronics and sometimes mains voltage — follow the creator’s safety notes.
More builds like this
All projectsGallery
Start here
Navigation into the creator’s own docs — we don’t rewrite the guide, we route you to the source.
- 1.Read the repository README for the board pinout and display wiring (Start here to confirm your ESP32-S3 variant matches the memory requirements: 8MB PSRAM and 16MB flash.)
- 2.Download the model with scripts/fetch_model.sh, then flash model and firmware with scripts/deploy.sh (The README contains flashing instructions; the published models are linked on Hugging Face.)
- 3.Test with the TinyStories model first, then try the Barista model if you want narrow Q&A behaviour (Both models are under 15MB and documented in the repo.)
KNOWN ISSUES
- The published models write short simple stories or answer espresso questions only — they will not follow instructions, write code, or know general facts, and retraining on a different dataset is a separate project.
- Not all ESP32-S3 boards have 8MB of PSRAM; cheaper variants with 2MB will not run the 28.9M parameter model.
- The 9.5 tokens per second speed is end-to-end generation including display refresh — it will feel slower than server-based models, and longer prompts will stretch the wait.
- Model conversion from PyTorch requires Python and familiarity with quantization scripts if you want to train your own; the published models avoid this step.
- The display wiring is optional but without it you will only see output over serial, which is harder to demo.
- The Per-Layer Embeddings architecture means inference speed depends on flash read latency — a board with slower flash will run slower.
Can I fine-tune it on my own dataset?
Yes, the repository includes model conversion scripts and training is standard PyTorch, but you will need a machine with a GPU and familiarity with the TinyStories or similar datasets to get coherent output from a model this small.
Why is it only trained on TinyStories?
A 28.9 million parameter model does not have the capacity for general knowledge or instruction-following — TinyStories is a constrained dataset designed for small models, and it lets the architecture demonstrate inference speed and memory efficiency without overpromising capability.
What is the practical use case?
This is primarily a research and educational project showing how far on-device inference can be pushed on microcontroller-class hardware; practical applications are narrow-domain text generation where connectivity is unavailable or undesirable.
How does Per-Layer Embeddings compare to other quantization methods?
Per-Layer Embeddings keep only the active rows of the embedding table in fast memory instead of the whole table; the measured benefit over a same-core baseline that fits in SRAM is about 9.3 percent perplexity reduction, and the parameter count scales roughly 110 times larger than earlier independent ESP32 work.
Community builds
No community builds yet — be the first, we feature the best ones.
Discussion1
FROM THE COMPAREE TEAM
The ESP32-S3 runs a 28.9 million parameter model at close to 10 tokens per second with no server behind it. What narrow domain would you want to run entirely on-chip?
slvDev
slvDev published esp32-ai in July 2026 and has kept iterating on memory-constrained inference for microcontrollers, applying architectural ideas from Google's Gemma 3n to hardware with 512KB of SRAM. The same repo now also runs part of the fruit fly connectome on the chip.
DISCLAIMER
- Comparee is not the author of the projects featured here. All rights to each project belong to its creator — every page links to the original source, and we never host creators’ files.
- Information is provided without warranty and may become outdated as projects evolve. Prices are indicative bands only — always check the creator’s parts list for current costs.
- Building and operating any project is at your own responsibility. Protective equipment, safe workshop practice and compliance with local regulations are the builder’s responsibility.
CompareeTEAM2mo agoedited
Practical notes from our verification: the repository is active (created in July 2026, with pushes since), and the trained models are published on Hugging Face and linked in the README, including TinyStories for story generation and Barista for espresso questions. The single biggest hardware decision is the board: the README's reference chip is an ESP32-S3 with 512KB SRAM, 8MB PSRAM and 16MB flash, because most of the model lives in flash and the core sits in PSRAM, so check both numbers before buying. The Per-Layer Embeddings approach, borrowed from Google's Gemma 3n, is explained in the README with a memory layout breakdown. Setup is script-based: one script downloads the model and a separate deploy script compiles and flashes the model and firmware. Correction (4 October 2026): we re-checked this page line by line against the project's own repository, documentation and videos, and fixed errors in earlier versions.