Prerequisites
Combinations
Mix and match to fit your hardware budget.RAM figures are approximate peak usage. When running LLM + STT + TTS
simultaneously, add the individual figures together.
“Fully local voice” means the audio, not the conversation. A local voice
stack transcribes on your machine and then sends the text to whatever LLM you
configured. Pair it with Ollama (below) to keep the whole turn local.
Setup
Troubleshooting
”Connection refused” from Ollama
Make sureollama serve is running. By default it listens on
http://localhost:11434. Check with:
Slow first response
The first inference after pulling a model is slow because Ollama loads weights into memory. Subsequent calls are fast. You can pre-warm with:Out of memory
If your machine runs out of RAM:- Switch to a smaller model (
qwen2:7buses ~1 GB less thanllama3:8b). - Close other heavy applications.
- On Linux, increase swap:
sudo fallocate -l 8G /swapfile && sudo mkswap /swapfile && sudo swapon /swapfile.
STT not detecting speech
Ensure your microphone is accessible and the correct device is selected:Performance Tips
Choose models based on your hardware:- 8 GB RAM (no GPU):
qwen2:7b+faster-whisper base+ Piper. The LLM dominates the turn; the endpointing and TTS costs are documented in Local Voice. - 16 GB RAM (no GPU):
llama3:8b+faster-whisper small+ Piper. - Apple Silicon (M1+): Ollama uses Metal acceleration automatically.
llama3:8bruns at ~30 tokens/s on M2. - NVIDIA GPU (8 GB+ VRAM): Ollama detects CUDA automatically. Expect 40–80 tokens/s depending on model and GPU.
- Keep Ollama running between sessions to avoid cold-start latency.
- Use the
basemodel unless you need higher accuracy;smallis meaningfully slower for a modest accuracy gain. - On macOS, prefer
whispercppoverfaster-whisper. CTranslate2 has no Metal backend, so faster-whisper runs on CPU there while every log line claims local acceleration. - Piper TTS is CPU-only. Its throughput has not been measured in this repository, because its synthesis path has never run here; see Local Voice.
