A LANGUAGE MODEL RUNS THIS FISH TANK, AND IT FITS ON AN 8-DOLLAR CHIP
A virtual aquarium where a language model decides how every fish behaves, running entirely on an ESP32 with no network and no cloud.
AIDisplays
Built withESP32
- difficulty
- ●●●○○
- time
- an evening to flash, a weekend-plus to train your own model
- license
- MIT
- repo
- repo ACTIVE141 stars
●●●○○ · an evening to flash, a weekend-plus to train your own model · MIT · 141 stars · repo ACTIVE
WHAT YOU’LL NEED
Jump to section
COMPAREE VERDICT
This is an unusually candid tiny-LLM project. Strato Doumanis distilled a 26-billion-parameter Gemma teacher down to 14.3 million parameters and got it to run inference at 12 tokens per second on the second core of an ESP32-S3 while the fish tank renders at 25-30 fps on the first. Each fish describes its situation in one line of structured text — hunger 7, energy 5, stress 2, curiosity 8, bold 4, a friend nearby — and the model completes it with a goal and an urgency, like 'seek_food urgency 8'. A separate reflex layer turns that into actual steering. The model picks the teacher's goal 72 percent of the time, and the teacher agrees with itself 82 percent of the time, so there is drift — and when the model is unsure, the fish visibly hesitates by design, which is oddly compelling. The vocabulary is 54 tokens, a closed schema lexicon not English, which is why it fits. If you want to run the published model, you flash it from the browser in a couple of clicks. If you want to train your own fish personalities or tune the distillation, you will spend a weekend-plus in PyTorch and Ollama. Hardware is a Waveshare ESP32-S3 1.8-inch AMOLED touch board with 8 MB PSRAM, 16 MB flash, a 368×448 screen, IMU and real-time clock, around 35 dollars per the README. This is a teaching project in what edge AI actually looks like when you strip out the hype.
IN THE REPO
GOOD TO KNOW
- —The trained 7.56 MB model file is in the repo and there is a two-click browser installer.
- —Complete distillation pipeline published — Python, PyTorch, Ollama for the teacher, around 51,000 training situations.
- —PC simulator included so you can test without hardware.
- —No PCB or enclosure files; this is bare board or 3D-print-your-own territory.
- —Licence MIT, no restrictions.
- —Repo includes detailed notes on token vocabulary design and quantisation trade-offs.
Parts to buy
2 itemsFrom our check of the build. Exact quantities and part numbers are in the creator’s BOM.
Can I build this?
Build at your own risk. Projects involve tools, electronics and sometimes mains voltage — follow the creator’s safety notes.
Videos
The creator's own YouTube series documents the build — distilling the model, designing the schema and first boot on real hardware.
More builds like this
All projectsGallery
Start here
Navigation into the creator’s own docs — we don’t rewrite the guide, we route you to the source.
- 1.Read the README's 'The numbers' and 'How it's put together' sections to understand the token vocabulary and decision pipeline before you flash anything (The vocabulary is not English — it is 54 tokens of fish schema — and understanding that is half the project.)
- 2.Use the two-click browser installer to flash the published model and firmware to your Waveshare board (Requires Chrome or Edge; the model file is 7.56 MB and flashes in under a minute.)
- 3.Run the PC simulator first if you want to watch the decision loop without hardware (Runs the same C code as the firmware in an LVGL + SDL2 window; --narrate prints every decision.)
- 4.If you want to train your own model, start with the distillation pipeline in the repo and Ollama for the teacher (The shipped model was trained on 51,613 teacher-labelled situations; expect a weekend-plus to regenerate data, tune and re-quantise.)
KNOWN ISSUES
- The 8-dollar figure is the bare ESP32-S3 chip; the Waveshare board with screen, battery and PSRAM used here is 35 dollars — still cheap, but five times more than the hook suggests.
- The model vocabulary is 54 tokens of closed schema, not English — you cannot ask it questions or add new fish traits without retraining and re-quantising the entire model.
- Inference takes about 3.7 seconds per decision at 12 tokens per second, so fish goals update every few seconds, not in real-time — expect visible pauses.
- The student model picks the teacher's goal 72 percent of the time, and the teacher agrees with itself 82 percent of the time — there is drift, and fish behaviour will surprise you.
- No enclosure files; the bare board sits on your desk or you 3D-print your own, and the repo does not include STLs.
- If you want to train your own model, you need PyTorch, Ollama, a machine that can run a 26-billion-parameter teacher, and a weekend to understand why the token vocabulary is that shape.
Can I change the fish behaviour without retraining the model?
No. The model was trained on 51,613 teacher-labelled situations built from a fixed state line — five drives (hunger, energy, stress, curiosity, boredom), three personality traits (bold, sociable, lazy), trust and life stage — and its 54-token vocabulary is closed. Adding a new trait or goal means regenerating data, retraining the student and re-quantising.
Why is inference so slow — about 3.7 seconds per decision?
The ESP32-S3 is running a 14.3-million-parameter transformer at 12 tokens per second, using the chip's SIMD instructions with weights read straight from flash. That is fast for a microcontroller; the tank renders at 25-30 fps while the model runs on the second core.
What happens if I use a bare ESP32-S3 instead of the Waveshare board?
You save 27 dollars but you lose the AMOLED screen, the battery, the IMU, the real-time clock, and the touch interface — you would need to add your own display and power, which is more work than the repo assumes.
Can I run this on a different ESP32 board?
Only if it has 8 MB PSRAM and 16 MB flash — the model file is 7.56 MB and the firmware expects that memory layout. Most ESP32 boards do not ship with enough PSRAM.
Community builds
No community builds yet — be the first, we feature the best ones.
Discussion1
FROM THE COMPAREE TEAM
The small model picks the teacher's goal 72 percent of the time, and the teacher agrees with itself 82 percent of the time, so fish sometimes hesitate. If you trained your own model, what fish personality would you distill first?
Strato Doumanis
Strato Doumanis (mediacutlet on GitHub) built Pocket Tank to show that edge AI does not need gigabytes of VRAM or a cloud API: a 14-million-parameter model distilled from a 26-billion-parameter teacher makes every fish's decisions at 12 tokens per second on an ESP32-S3. By design the model owns its decisions, and when it is unsure the fish visibly hesitates instead of being overridden by a hidden fallback.
DISCLAIMER
- Comparee is not the author of the projects featured here. All rights to each project belong to its creator — every page links to the original source, and we never host creators’ files.
- Information is provided without warranty and may become outdated as projects evolve. Prices are indicative bands only — always check the creator’s parts list for current costs.
- Building and operating any project is at your own responsibility. Protective equipment, safe workshop practice and compliance with local regulations are the builder’s responsibility.
CompareeTEAM23d agoedited
Practical notes from our verification: the repo includes the trained model, the distillation pipeline, a PC simulator and a browser installer, and the licence is MIT. The cheap chip in the headline is the bare ESP32-S3; the README itself says the board used here, a Waveshare ESP32-S3 with a 1.8-inch AMOLED screen and battery, costs around 35 dollars. A decision takes about 3.7 seconds on the real board, so fish goals update every few seconds rather than instantly, and the visible hesitation is part of the design. The model's vocabulary is 54 tokens of a closed schema, not English, so new fish traits mean retraining. If you just want to run it, you flash from the browser and you are done. If you want to train your own fish personalities, the pipeline is documented, but the training data was labelled by a 26-billion-parameter teacher model, so expect a capable machine and a weekend or more. Correction (4 October 2026): we re-checked this page line by line against the project's own repository, documentation and videos, and fixed errors in earlier versions.