BUILD A TALKING AI ASSISTANT ON A BREADBOARD AND HOST ITS BRAIN YOURSELF
A breadboard, five cheap parts, and it talks back to you — with the wake word running on the chip itself.
by 78
AIOpen-hardware
Built withESP32
- difficulty
- ●●●○○
- time
- a weekend
- license
- MIT
- repo
- repo ACTIVE30,399 stars
●●●○○ · a weekend · MIT · 30,399 stars · repo ACTIVE
WHAT YOU’LL NEED
Jump to section
COMPAREE VERDICT
XiaoZhi is an open-source voice assistant firmware for ESP32 boards, and you can build one yourself on a breadboard from cheap modules (an ESP32-S3, an I2S microphone, an I2S amplifier with a small speaker and an optional OLED), following the wiring tutorial linked from the README, or flash one of 171 prebuilt variants for supported boards. Wake-word detection runs offline on the chip with Espressif's ESP-SR and is customisable, so the device is not streaming audio until you address it. What makes it more than a toy chatbot is MCP: device-side MCP lets the model control the hardware it lives in (speaker, LED, servo, GPIO), and cloud-side MCP extends it to smart home control and more. The honest limitation is that the device needs a network connection and a server to think: by default it connects to the official xiaozhi.me server, where personal users can use the Qwen real-time model for free. If you want to run the server side yourself, the README lists community servers such as xinnan-tech/xiaozhi-esp32-server (Python, MIT). The biggest trap is mistaking the on-chip wake word for a fully offline assistant. If you are happy with the official server, this is one of the cleanest ESP32 voice builds available.
IN THE REPO
GOOD TO KNOW
- —Prebuilt firmware binaries for 171 variants across 138 board directories — no compilation required for the reference build.
- —Full wiring diagrams and pin assignments in the docs; no custom PCB, jumper wires are enough.
- —MIT license, no commercial restrictions.
- —The device needs a network connection and a backend to think — out of the box it pairs with a hosted console.
- —Community self-hosted servers exist if you do not want to use the official one, for example xinnan-tech/xiaozhi-esp32-server (Python, MIT), plus Java and Go alternatives listed in the README.
- —Wake-word detection runs offline on the chip using Espressif's ESP-SR, with customizable wake words — audio is not streamed until you address it.
Parts to buy
6 itemsFrom our check of the build. Exact quantities and part numbers are in the creator’s BOM.
Can I build this?
Build at your own risk. Projects involve tools, electronics and sometimes mains voltage — follow the creator’s safety notes.
Videos
I Built an AI Desk Buddy with ESP32 (Xiaozhi + Custom Face UI) - Tech Talkies
The project's own videos are on Bilibili (in Chinese), linked at the top of the README, including a beginner's build guide; this English video from Tech Talkies shows a community build with a custom face UI.
More builds like this
All projectsGallery
Start here
Navigation into the creator’s own docs — we don’t rewrite the guide, we route you to the source.
- 1.Pick your board(The breadboard DIY build uses a generic ESP32-S3 dev board (the bread-compact-wifi target), but 138 board directories are supported — check the firmware matrix in the repo for yours.)
- 2.Follow the wiring diagram(Pin assignments are in the docs; the reference build uses jumper wires, no soldering.)
- 3.Flash the prebuilt firmware(171 firmware variants are available — download the binary for your board and flash it.)
- 4.Configure network and backend(Out of the box it pairs with the hosted console; if you want self-hosted, set up xiaozhi-esp32-server separately.)
Resources
Documentation, files and community threads for this build — we link straight to the original sources and never rehost the creator’s files.
KNOWN ISSUES
- The on-chip wake word is not the same as a fully offline assistant — only wake-word detection runs locally, the reasoning still needs a server.
- If you do not want to use the hosted backend, you will need to run xiaozhi-esp32-server yourself, which is a separate setup step.
- The firmware matrix covers 138 boards, but if yours is not listed you will be compiling from source.
- Interrupting it mid-sentence (realtime full-duplex) needs AEC-capable hardware; a simple breadboard build may not support it, so check your board.
- MCP integration is powerful but assumes you are comfortable with the hosted console or willing to configure the self-hosted backend to reach it.
- The display is optional, but the emoji expressions on an OLED or LCD are a big part of the charm.
Is this fully offline?
No. Wake-word detection runs on the chip itself using Espressif's ESP-SR, so it is not streaming audio until you address it, but the reasoning and response generation need a network connection and a backend — either the hosted console or the self-hosted server.
Do I need to compile the firmware?
Not for the reference build or any of the 171 supported variants — prebuilt binaries are available. If your board is not in the list, you will compile from source.
What is the difference between the hosted backend and the self-hosted one?
The hosted console is ready to use out of the box; the self-hosted backend (xinnan-tech/xiaozhi-esp32-server) runs on your own machine and gives you full control, but it is a separate setup step.
Can it control smart home devices?
Yes, through cloud-side MCP — but that assumes you are using a backend that supports MCP integration, either the hosted one or your own.
Community builds
No community builds yet — be the first, we feature the best ones.
Discussion1
FROM THE COMPAREE TEAM
Over 30,000 stars, 171 firmware variants, and the wake word runs on the chip itself — but the conversation still needs a server. Would you run your own backend, or is the free official server good enough for a weekend build?
78
78 built XiaoZhi as an open-source, MCP-based ESP32 voice assistant with on-chip wake-word detection and released the firmware under the MIT licence. It has grown to over 30,000 stars, and the community has built several self-hostable servers for it.
DISCLAIMER
- Comparee is not the author of the projects featured here. All rights to each project belong to its creator — every page links to the original source, and we never host creators’ files.
- Information is provided without warranty and may become outdated as projects evolve. Prices are indicative bands only — always check the creator’s parts list for current costs.
- Building and operating any project is at your own responsibility. Protective equipment, safe workshop practice and compliance with local regulations are the builder’s responsibility.
CompareeTEAM2mo agoedited
Practical notes from our verification: the repo is actively developed and supports 138 board directories and 171 release variants, and the README also shows a breadboard DIY build with a wiring tutorial if you want to start from loose modules. The single biggest decision is the backend: by default the firmware connects to the official xiaozhi.me server, where personal users can register and use the Qwen real-time model for free, or you can run your own server using one of the community projects the README lists, such as xinnan-tech/xiaozhi-esp32-server in Python. The offline wake word using Espressif's ESP-SR is real and customisable, but the speech recognition, model and voice responses still need a network connection — this is not a fully offline assistant. The MCP integration is the standout feature: device-side MCP lets the model control the hardware it lives in (speaker, LED, servo, GPIO), and cloud-side MCP extends it to smart home control and more. Beginners can flash ready-made firmware without setting up a development environment, and OLED and LCD displays get emoji expressions. Correction (4 October 2026): we re-checked this page line by line against the project's own repository, documentation and videos, and fixed errors in earlier versions.