voicemate logo: a cyberpunk assistant with a headset and a voice waveform on her visor

voicemate

A private voice assistant that runs entirely on your Mac,
in English and Hungarian

GitHub Docs

Why I built it

I wanted an assistant I can simply talk to, in Hungarian or English, the way I would talk to a colleague. Cloud assistants send conversations to third parties, forget context between sessions and handle Hungarian poorly. My first prototype was a blocking five-second record, transcribe, generate, speak loop: Hungarian only, no tools, no memory and too slow to feel like a conversation.

voicemate is the rewrite. Speech recognition, the language model, memory and speech synthesis all run on the machine. Only explicit tool calls (web search, page fetch, arXiv, weather) go online, and each one is marked in the UI. The assistant's name ("Ava" by default) is just a setting.

Early look

voicemate in the browser: a conversation in the middle, pipeline stages with timings on the left, recalled memory on the right

A real session with the fast profile (gemma4:e4b): a tool-backed question about the date and weather, then a trip-planning request interrupted mid-answer with a follow-up. The left rail shows each pipeline stage with its timing, the right panel what was recalled from memory.

How it works

Each part is shown as the diagram it is built from. Hover a diagram to magnify it, click to open it full size.

plain code speech model language model / agent safeguard / goes online input / output / storage

Streaming pipeline speech

Everything runs on the Mac. Speech starts while the answer is still being written, so a chat turn typically begins speaking 1-3 s after you stop talking.

  • Silero VAD finds the end of your turn; Parakeet v3 transcribes it; the language is detected per turn and the reply comes back in the same language.
  • Piper speaks Hungarian and Kokoro speaks English, one sentence at a time.
  • The speech models run on the CPU on purpose: the GPU and most of the memory belong to the language model.

Agent and tools LangGraph agent

A small graph: recall loads relevant memory, the agent calls the model with every tool bound, and tools run what it asks for.

  • Tools: web search, page fetch, arXiv search and paper download, weather, time, calculator, notes and reminders, and files in two whitelisted folders.
  • Errors go back to the model as text so it can retry; after 6 tool rounds the tools are removed and it must answer.
  • A SQLite checkpointer keeps the conversation across page reloads and restarts.

Natural turn-taking interruptions

Speak while it talks and it stops at once, remembers what you already heard, and treats your words as a correction or a new question.

  • Barge-in needs louder and longer speech than a normal turn, so its own voice from the speakers does not interrupt it.
  • Hold lets you think out loud: pauses no longer end the turn, and everything is answered at once when you say you are finished.
  • Ctrl+M turns the microphone on and off; turning it off means "I am done" and what was heard is answered.

One spoken turn latency

The latency budget of a chat turn, from the last word you say to the first audio you hear.

  • The prompt prefix never changes between turns, so Ollama reuses its cached prompt; only the newest message carries the time, language and recalled memories.
  • A tool that goes online adds a short spoken filler such as "One moment, let me check."

Under the hood

Local & private
  • Conversations, memory, notes and logs stay in a git-ignored folder; raw audio is not stored.
  • The server only answers to localhost, and fetched URLs must resolve to public addresses.
  • Web content is treated as untrusted data, never as instructions.
Memory
  • LanceDB with bge-m3 embeddings holds facts about me, past conversations and a research cache.
  • Asking the same research question again is answered from memory, with the date it was found.
Asks before risky actions
  • Replacing an existing file, or opening a page no search result listed, pauses the agent with LangGraph's interrupt().
  • The question is shown and spoken, and your exact words resume the run, so a "no, call it plan2" redirects the action.
Measured, not guessed
  • Parakeet transcribes a 4 s utterance in about 280 ms on the CPU; Piper speaks a sentence in about 70 ms.
  • Three model profiles (fast, gemma, qwen); a startup memory check falls back to fast if the chosen one would be paged out.
  • make bench and make report give component benchmarks and a latency summary of recent turns.

A personal, local-first project.


voicemate is developed on an Apple Silicon Mac (M4, 32 GB) with Ollama, needs about 15 GB of disk for the models and is not packaged on PyPI: you run it from a clone of the repository. The code and the diagrams above are the current state.


Stack: Python, LangGraph, Ollama, NiceGUI, LanceDB, Silero VAD, Parakeet (MLX), Piper, Kokoro. MIT licensed.

Star it on GitHub