A 28.9M parameter language model runs entirely on a cheap ESP32-S3 microcontroller

A 28.9M PARAMETER LANGUAGE MODEL RUNS ENTIRELY ON A CHEAP ESP32-S3 MICROCONTROLLER

A 28.9 million parameter model runs entirely on a cheap ESP32-S3 microcontroller with no server behind it.

by slvDev

FULL CAD BOM FIRMWARE DOCS

AIOpen-hardware

Built withESP32Arduino

difficulty
●●●○○
time
a weekend
license
MIT
repo
repo ACTIVE4,458 stars
1
Jump to section

COMPAREE VERDICT

This is a striking demonstration of a twenty-nine-million parameter language model running entirely on a microcontroller with no connectivity. The ESP32-S3 has 512KB of SRAM and 8MB of PSRAM, so the model should not fit — but slvDev borrowed Per-Layer Embeddings from Google's Gemma 3n and parked 25 million parameters in a lookup table in slow flash. Each token reads roughly 450 bytes out, activations stay in SRAM, and the dense core sits in PSRAM. The whole 28.9M parameter model quantized to 4 bits is 14.9MB. It generates at roughly 9.5 to 9.9 tokens per second end to end, displayed on a small screen. The measured benefit over a same-core baseline that fits in SRAM is 0.098 nats, about 9.3 percent perplexity improvement. For scale, earlier independent work on the same chip family ran 260 thousand parameters, so this is roughly 110 times larger. The honest limit is that the published language models were trained on TinyStories or a narrow espresso Q&A dataset — they will not answer general questions, follow instructions, write code, or know facts. If you want to understand how small models are deployed on constrained hardware or explore architectural tricks at the memory boundary, this is a very clean example. If you want a general-purpose assistant, this is not it and training your own will take more than a weekend. The single thing most likely to go wrong is expecting ChatGPT behaviour from a 28.9M parameter TinyStories model.

GOOD TO KNOW

  • —MIT licence permits commercial use.
  • —Repository contains firmware, training and export code, and three published models on Hugging Face (TinyStories, Barista and the fruit fly connectome demo).
  • —No formal bill of materials: the README names the chip (ESP32-S3 N16R8: 8MB PSRAM, 16MB flash) and the firmware drives a small SH110X OLED display.
  • —No CAD files — this is a software project with off-the-shelf hardware.
  • —Documentation covers flashing, model conversion from PyTorch, and the Per-Layer Embeddings architecture.
  • —The trained models are TinyStories-based: they write short simple stories, not general-purpose text.

Parts to buy

3 items

From our check of the build. Exact quantities and part numbers are in the creator’s BOM.

  • ESP32-S3 dev board with 16MB flash and 8MB PSRAM (N16R8)Find
  • Optional small SH110X OLED displayFind
  • USB cableFind

BUILDS OF THE WEEK

Five open-source builds worth your weekend, every week.

Checked like this one: what’s really in the repo, what it costs, how hard it is. One email, unsubscribe anytime.

Can I build this?

Printnothing required
BuyESP32-S3 dev board with 16MB flash and 8MB PSRAM (N16R8), optional small SH110X OLED display, USB cable
Toolsarduino-cli (used by the deploy script), Python with uv, and a GPU machine only if training your own model
SkillsComfortable flashing firmware to microcontrollers and reading Python scripts; no ML experience required to run the published models
TimeTwo to four hours to flash and test the published models, add a weekend if training your own on a different dataset
Cost$ — dominated by the ESP32-S3 board, everything else is a USB cable and optional display under ten dollars total
SafetyNone beyond ordinary electronics care — USB-powered, no mains voltage, no lithium cells to charge.

Build at your own risk. Projects involve tools, electronics and sometimes mains voltage — follow the creator’s safety notes.

More builds like this

All projects

Gallery

github.com
github.com
github.com

Start here

Navigation into the creator’s own docs — we don’t rewrite the guide, we route you to the source.

  1. 1.Read the repository README for the board pinout and display wiring (Start here to confirm your ESP32-S3 variant matches the memory requirements: 8MB PSRAM and 16MB flash.)
  2. 2.Download the model with scripts/fetch_model.sh, then flash model and firmware with scripts/deploy.sh (The README contains flashing instructions; the published models are linked on Hugging Face.)
  3. 3.Test with the TinyStories model first, then try the Barista model if you want narrow Q&A behaviour (Both models are under 15MB and documented in the repo.)

KNOWN ISSUES

  • The published models write short simple stories or answer espresso questions only — they will not follow instructions, write code, or know general facts, and retraining on a different dataset is a separate project.
  • Not all ESP32-S3 boards have 8MB of PSRAM; cheaper variants with 2MB will not run the 28.9M parameter model.
  • The 9.5 tokens per second speed is end-to-end generation including display refresh — it will feel slower than server-based models, and longer prompts will stretch the wait.
  • Model conversion from PyTorch requires Python and familiarity with quantization scripts if you want to train your own; the published models avoid this step.
  • The display wiring is optional but without it you will only see output over serial, which is harder to demo.
  • The Per-Layer Embeddings architecture means inference speed depends on flash read latency — a board with slower flash will run slower.

Can I fine-tune it on my own dataset?

Yes, the repository includes model conversion scripts and training is standard PyTorch, but you will need a machine with a GPU and familiarity with the TinyStories or similar datasets to get coherent output from a model this small.

Why is it only trained on TinyStories?

A 28.9 million parameter model does not have the capacity for general knowledge or instruction-following — TinyStories is a constrained dataset designed for small models, and it lets the architecture demonstrate inference speed and memory efficiency without overpromising capability.

What is the practical use case?

This is primarily a research and educational project showing how far on-device inference can be pushed on microcontroller-class hardware; practical applications are narrow-domain text generation where connectivity is unavailable or undesirable.

How does Per-Layer Embeddings compare to other quantization methods?

Per-Layer Embeddings keep only the active rows of the embedding table in fast memory instead of the whole table; the measured benefit over a same-core baseline that fits in SRAM is about 9.3 percent perplexity reduction, and the parameter count scales roughly 110 times larger than earlier independent ESP32 work.

Community builds

No community builds yet — be the first, we feature the best ones.

Discussion1

FROM THE COMPAREE TEAM

The ESP32-S3 runs a 28.9 million parameter model at close to 10 tokens per second with no server behind it. What narrow domain would you want to run entirely on-chip?

CompareeTEAM2mo agoedited

Practical notes from our verification: the repository is active (created in July 2026, with pushes since), and the trained models are published on Hugging Face and linked in the README, including TinyStories for story generation and Barista for espresso questions. The single biggest hardware decision is the board: the README's reference chip is an ESP32-S3 with 512KB SRAM, 8MB PSRAM and 16MB flash, because most of the model lives in flash and the core sits in PSRAM, so check both numbers before buying. The Per-Layer Embeddings approach, borrowed from Google's Gemma 3n, is explained in the README with a memory layout breakdown. Setup is script-based: one script downloads the model and a separate deploy script compiles and flashes the model and firmware. Correction (4 October 2026): we re-checked this page line by line against the project's own repository, documentation and videos, and fixed errors in earlier versions.

slvDev

slvDev published esp32-ai in July 2026 and has kept iterating on memory-constrained inference for microcontrollers, applying architectural ideas from Google's Gemma 3n to hardware with 512KB of SRAM. The same repo now also runs part of the fruit fly connectome on the chip.

GitHub

Star the project on GitHub

DISCLAIMER

  • Comparee is not the author of the projects featured here. All rights to each project belong to its creator — every page links to the original source, and we never host creators’ files.
  • Information is provided without warranty and may become outdated as projects evolve. Prices are indicative bands only — always check the creator’s parts list for current costs.
  • Building and operating any project is at your own responsibility. Protective equipment, safe workshop practice and compliance with local regulations are the builder’s responsibility.