THIS LITTLE BOX RUNS A VOICE-CONTROLLED AI AGENT FULLY OFFLINE ON ONE RASPBERRY PI

A tiny Raspberry Pi box that listens for your voice and runs an LLM agent entirely locally - no cloud AI provider - and can use tools to act on your Pi.

by Simone (syxanash)

FULL CAD BOM FIRMWARE DOCS

AIOpen-hardware

Built withRaspberry Pi

difficulty
●●●●○
time
a weekend-plus
license
GPL-3.0
repo
repo ACTIVE354 stars
1
Jump to section

COMPAREE VERDICT

Max Headbox is a fully local voice-activated LLM agent for a Raspberry Pi 5: Vosk listens for the wake word, faster-whisper transcribes you, and small Gemma models running in Ollama decide what to do, with an animated character on a small screen showing what it is thinking. Its party piece is tools — small JavaScript modules it can call to check the weather, take notes or control hardware, and any tool you flag as dangerous asks you for a YES/NO confirmation first. Expect responses measured in seconds, not milliseconds; for a private agent that never phones home, that is the trade-off.

GOOD TO KNOW

  • —Complete software stack in the repo plus a short hardware list (Pi 5, microphone, optional GeeekPi screen case).
  • —The README covers installation (Node, Python backend, Ollama models), configuration of the .env file, running the app and writing your own tools.
  • —No custom electronics — just a Pi 5, a USB microphone and an optional screen case.
  • —GPL-3.0 license permits commercial use with source disclosure.
  • —The creator built this as a learning project and documents the entire reasoning behind model choices.
  • —Actively developed - last updated in August 2026, with the backend rewritten in Express.js.

Parts to buy

2 items

From our check of the build. Exact quantities and part numbers are in the creator’s BOM.

  • USB microphoneFind
  • Active cooler is a mustFind

BUILDS OF THE WEEK

Five open-source builds worth your weekend, every week.

Checked like this one: what’s really in the repo, what it costs, how hard it is. One email, unsubscribe anytime.

Can I build this?

Printnothing — the box is an off-the-shelf GeeekPi screen case
BuyRaspberry Pi 5 (8GB or 16GB, as tested by the creator), a USB microphone, and optionally the GeeekPi screen + case + active cooler bundle; an active cooler is a must.
ToolsA screwdriver for the case, and SSH access to configure the Pi (or a monitor and keyboard for first boot). No soldering.
SkillsIntermediate - comfort with the Linux command line, Node and Python package installs, systemd and editing a .env file. Writing tools needs basic JavaScript.
TimeAn evening or two: mount the Pi in the case (if you use the bundle), install Node, Python and Ollama, pull the two Gemma models, configure the .env and test the wake word and tools.
Costan 8 GB or 16 GB Raspberry Pi 5 is most of the cost, plus a USB microphone and optionally the GeeekPi screen, case and cooler bundle. No other parts are needed.
SafetyOrdinary low-voltage electronics. Use an active cooler, since sustained LLM inference heats the Pi 5. Tools you mark as non-dangerous run without asking, so flag anything that changes your system with dangerous: true.

Build at your own risk. Projects involve tools, electronics and sometimes mains voltage — follow the creator’s safety notes.

Videos

Max Headbox an LLM agent that runs on a Raspberry Pi

The creator's own demo, showing the agent asking for confirmation before running a tool.

More builds like this

All projects

Gallery

opengraph.githubassets.com
radar.comparee.ai
techcrunch.com
img.youtube.com
img.youtube.com

Start here

Navigation into the creator’s own docs — we don’t rewrite the guide, we route you to the source.

  1. 1.Read the architecture overview (Understand the pipeline - Vosk listens for the wake word, faster-whisper transcribes you, Gemma models in Ollama run the agent, and the web UI shows an animated face and the answer - before installing anything.)
  2. 2.Get the hardware: a Raspberry Pi 5, a USB microphone and (optionally) the GeeekPi screen, case and cooler bundle(The bundle is optional, but the creator insists on an active cooler. You can run it in any form factor as long as about 6 GB is free for the models.)
  3. 3.Gather the BOM (The creator tested on 8 GB and 16 GB Pi 5 boards and links the exact microphone he used.)
  4. 4.Install the software stack (Setup is manual: install Node 22, Python 3 and Ollama, run npm install and pip3 install, pull gemma3:1b and gemma4:e2b, expose Ollama on the network and fill in the .env addresses.)

Resources

Documentation, files and community threads for this build — we link straight to the original sources and never rehost the creator’s files.

KNOWN ISSUES

  • The creator tested on 8GB and 16GB Pi 5 boards and says to leave about 6GB free for the models — a 4GB Pi will struggle.
  • The wake word runs on Vosk's small English model, so it is tuned for English and a reasonably quiet room; a cheap or distant microphone will hurt both wake-word detection and transcription.
  • Skipping active cooling — the creator insists on an active cooler, because sustained LLM inference will throttle a bare Pi 5.
  • You need the models downloaded before the voice agent works — the README pulls gemma3:1b and gemma4:e2b through Ollama, and you should have about 6GB available to run them.
  • Response latency is not a bug - small local models on a Pi take seconds, not milliseconds, to answer. If you need instant replies, this hardware cannot deliver it.
  • Tools run without asking unless you mark them dangerous: true — do that for anything that changes your system, so the agent asks YES/NO first.

Can I use a Raspberry Pi 4?

It is not tested on a Pi 4 - the creator used 8 GB and 16 GB Pi 5 boards and says you need about 6 GB free for the models. An 8 GB Pi 4 might run it, but expect it to be noticeably slower.

Does it work in languages other than English?

Out of the box it is set up for English (the bundled Vosk wake-word model is en-US). faster-whisper and Gemma handle other languages, so you would need to swap the Vosk model and adjust prompts yourself.

How much does response time improve with a bigger model?

It does not — bigger models are slower. The README pulls two small Gemma models (Gemma 3 1B and Gemma 4 E2B) through Ollama; a larger model trades speed for coherence, and on a Pi the speed cost is brutal.

Can I add a camera or make it control smart home devices?

Yes, that is what the tool system is for: you add a JavaScript module in src/tools/ with a name, parameters, a description and an execute function, plus an Express route in backend/notions/ if it needs Pi hardware. The bundled tools (notes, weather, Wikipedia, CPU temperature, restart and more) show the pattern; mark anything risky as dangerous so it asks first.

Community builds

No community builds yet — be the first, we feature the best ones.

Discussion1

FROM THE COMPAREE TEAM

Over 350 stars and the speech and language models all run on a single Raspberry Pi, with no cloud AI provider involved. Would you trade response speed for that level of privacy?

CompareeTEAM2mo agoedited

Practical notes from our verification: the repository is actively maintained, and the README links two official demo videos from the creator, including one showing the confirmation flow for tools marked as dangerous. Wake-word detection uses Vosk with a small English model, transcription uses faster-whisper, and the language model runs locally through Ollama (gemma3:1b or gemma4:e2b), so expect to need around 6 GB of free memory for the LLMs on a Raspberry Pi 5. The creator is candid about trade-offs: this is not a snappy assistant, it is a working proof that a fully local voice agent is viable. The tool layer is the genuinely interesting piece — tools are JavaScript modules you add yourself, and the README explains why it avoids model tool-call APIs. Correction (4 October 2026): we re-checked this page line by line against the project's own repository, documentation and videos, and fixed errors in earlier versions.

Simone (syxanash)

Simone (syxanash) built Max Headbox as a fully local voice agent on a Raspberry Pi and wrote up the journey on his blog, including why he defines his own tool-calling format instead of using model tool-call APIs. The README's FAQ is refreshingly candid about shortcuts and planned changes.

GitHub

Star the project on GitHub

DISCLAIMER

  • Comparee is not the author of the projects featured here. All rights to each project belong to its creator — every page links to the original source, and we never host creators’ files.
  • Information is provided without warranty and may become outdated as projects evolve. Prices are indicative bands only — always check the creator’s parts list for current costs.
  • Building and operating any project is at your own responsibility. Protective equipment, safe workshop practice and compliance with local regulations are the builder’s responsibility.