TURN ANY PLUSHIE INTO A TALKING AI FRIEND WITH AN ESP32 CHIP INSIDE
An ESP32-S3 board running realtime voice AI from OpenAI, Gemini, ElevenLabs and others turns any stuffed toy into a voice AI that responds in real time — the maker open-sourced his entire startup stack.
by akdeb
AIHome
- difficulty
- ●●●●○
- time
- a weekend-plus
- license
- MIT
- repo
- repo ACTIVE2,029 stars
●●●●○ · a weekend-plus · MIT · 2,029 stars · repo ACTIVE
WHAT YOU’LL NEED
Jump to section
COMPAREE VERDICT
ElatoAI is a complete voice AI toy platform that streams realtime speech models — OpenAI Realtime, Gemini Live, xAI Grok, ElevenLabs, Hume and more — to an ESP32-S3 board, giving any plushie or toy real-time conversation with custom characters and voices from more than 100 supported models. The maker open-sourced the entire stack — firmware, web app, and backend — after running it as a product. Everything you need for the electronics side is documented, the firmware builds with PlatformIO or Arduino IDE, and the web app is Next.js. What is missing is the toy part: there is no enclosure STL, no guide to cutting a seam in a plushie and hiding the board inside, no battery holder design. You are expected to solve that. The real friction is not the hardware — it is the server and API setup. The quickest path registers your device with Elato's hosted server; if you self-host, you need an API key for at least one voice provider, a Supabase project, and an edge server on Deno or Cloudflare Workers, and every conversation is billed by the provider. The author does not quote a per-hour figure, and realtime voice APIs are not cheap; this is not a set-and-forget project. If you have built ESP32 audio projects before and you are comfortable managing a live API backend, this is a remarkably complete handoff of a working product. If you have never deployed edge functions or a database, expect an extra weekend just getting your own infrastructure running. The one thing most likely to go wrong is underestimating the ongoing cost of the paid voice APIs.
IN THE REPO
GOOD TO KNOW
- —MIT license, full commercial use allowed.
- —Complete firmware for ESP32-S3, a Next.js web app, edge server code (Deno or Cloudflare Workers, plus a FastAPI option) and a Supabase schema are all present.
- —Hardware BOM specified: ESP32-S3 module (no PSRAM required), INMP441 microphone, MAX98357A amplifier, speaker.
- —Setup requires an API account with a voice AI provider (OpenAI, Gemini, xAI, ElevenLabs, Hume or Boson) and ongoing per-use costs — this is not a one-time purchase.
- —Documentation covers board setup, server deployment and the web app, and assumes you are comfortable with PlatformIO or Arduino IDE, npm and Supabase.
- —No 3D-printable enclosure or plushie integration guide — you design that part yourself.
Parts to buy
4 itemsFrom our check of the build. Exact quantities and part numbers are in the creator’s BOM.
Can I build this?
Build at your own risk. Projects involve tools, electronics and sometimes mains voltage — follow the creator’s safety notes.
Videos
Build Conversational AI Voice Clones with ElevenLabs on ESP32 Arduino | Elato
More builds like this
All projectsGallery
Start here
Navigation into the creator’s own docs — we don’t rewrite the guide, we route you to the source.
- 1.Read the full README (covers all three components: firmware, server, app)
- 2.Order the ESP32-S3 module and audio components(exact part numbers listed; any ESP32-S3 works, PSRAM is not required)
- 3.Set up OpenAI API account and budget alerts(Realtime API pricing is per-token; set a spending cap before you test)
- 4.Deploy the relay server (Deno Edge Functions or Cloudflare Workers (a FastAPI version is also included); the database is Supabase)
- 5.Flash firmware via PlatformIO(set the server URL before you flash; connect the device to WiFi afterwards through its captive portal. Arduino IDE is also supported.)
- 6.Deploy the Next.js web app(the web app (Next.js, deployable to Vercel) is where you create characters, manage devices and control volume from your phone's browser)
Resources
Documentation, files and community threads for this build — we link straight to the original sources and never rehost the creator’s files.
KNOWN ISSUES
- OpenAI Realtime API costs per token, and there is no per-conversation estimate in the docs — set a spending cap on your API key before you test or you will be surprised.
- Buy an ESP32-S3 board, not a classic ESP32 or C3 — the firmware targets the S3. PSRAM is not required.
- The server address is set in the firmware before you flash, so deploy the server first; WiFi is configured afterwards through the device's captive portal, and later firmware updates can go over the air.
- There is no enclosure design or plushie integration guide — you are expected to figure out how to hide the board, speaker, and battery inside a toy without a hot-glue disaster.
- The relay server is a single point of failure — if it goes down or rate-limits, every toy stops working; plan for monitoring and restarts.
- Battery size and runtime are not specified — you pick the capacity, but more capacity means more weight and bulk inside the toy.
Does this work offline?
Not with the main ElatoAI stack: every conversation streams through an edge server to a cloud voice AI provider, so it needs internet. Since March 2026 the author also publishes a separate Local AI Toys project (akdeb/local-ai-toys) that runs the conversation on local models instead.
Can I use a different LLM?
Yes. Out of the box the server supports OpenAI Realtime, Gemini Live, xAI Grok Voice, ElevenLabs Conversational AI agents, Hume EVI-4 and Boson Higgs Realtime; you pick the provider per character and supply that provider's API key.
How much does each conversation cost?
The README does not give a per-conversation figure. OpenAI charges per audio token, and Realtime API pricing is higher than text. Set a spending cap and monitor usage for your first few days.
What battery should I use?
The open-source docs do not specify a battery — the board is powered over USB. If you want it cordless, you have to choose a LiPo pack and charging board yourself and measure the runtime.
Can I run multiple toys from one server?
Yes — the relay server is designed to handle multiple concurrent websocket connections, one per toy. Each toy gets its own OpenAI session.
Community builds
No community builds yet — be the first, we feature the best ones.
Discussion1
FROM THE COMPAREE TEAM
The code is MIT-licensed and the docs name every component and pin, but there is no enclosure file and no battery-life estimate. What would you hide yours inside first?
akdeb
The creator built ElatoAI as a startup product, then open-sourced the entire stack — firmware, app, and server — after deciding to release it. The project reached the front page of Hacker News as a Show HN.
DISCLAIMER
- Comparee is not the author of the projects featured here. All rights to each project belong to its creator — every page links to the original source, and we never host creators’ files.
- Information is provided without warranty and may become outdated as projects evolve. Prices are indicative bands only — always check the creator’s parts list for current costs.
- Building and operating any project is at your own responsibility. Protective equipment, safe workshop practice and compliance with local regulations are the builder’s responsibility.
CompareeTEAM2mo agoedited
Practical notes from our verification: the repository is active and splits cleanly into ESP32 firmware, server code and a Next.js web app, with Supabase as the database; most of the step-by-step documentation lives on elatoai.com/docs rather than in the repo READMEs. The docs list the parts (ESP32-S3, INMP441 microphone, MAX98357A amplifier, speaker, LED and a button or touchpad) with exact pin connections, and the board is powered over USB, so battery choice is left to you and there is no enclosure model to print. You pick the voice back end yourself: OpenAI Realtime, Gemini Live, Grok, ElevenLabs, Hume or Boson through the Deno server, or other models through Cloudflare Workers or a FastAPI server. The README gives no estimate of cost per conversation, and these providers bill by usage, so keep an eye on your account. If you have deployed websocket servers before, this is a remarkably complete handoff; if not, budget an extra weekend for the server side. Correction (4 October 2026): we re-checked this page line by line against the project's own repository, documentation and videos, and fixed errors in earlier versions.