Local Voice Assistant for Smart Home
What this article covers
- How to build a completely local voice assistant.
- Whisper (STT) + Ollama (comprehension) + Piper (TTS).
- Connecting the assistant to Home Assistant.
- Practical examples for lights, climate, music, and security.
- Best practices for latency, accuracy, and privacy.
Introduction: Local voice assistants explained
A local voice assistant understands voice commands without the cloud. You speak, Whisper transcribes, Ollama understands and decides, Piper responds. Your data never leaves your network, unlike Alexa or Google Home.
This article is for users who want to build a local voice assistant. For fundamentals, check out Home Assistant with AI and Voice Assistants.
Why do you need a local voice assistant?
Alexa and Google Home send your voice commands to the cloud. A local assistant processes everything on your server. “Turn on the light” gets transcribed, understood, and executed locally. No data leaves, no ads, full control.
How a local voice assistant works
Microphone → Whisper (STT) → Ollama (LLM) → Home Assistant (action) → Piper (TTS) → Speaker. All local, no cloud.
The core idea: voice control without data transmission.
Who is this article for?
- Privacy-conscious users who don’t want cloud voice assistants.
- Smart home tinkerers building local voice control.
- Home Assistant users running Assist locally.
- Developers building custom voice assistants.
Key terms
- Whisper - OpenAI’s speech-to-text (local). Useful for: voice input.
- Ollama - Local model server. Useful for: comprehension.
- Piper - Text-to-speech (local). Useful for: voice output.
- Wyoming - Protocol for voice services. Useful for: integration.
- Home Assistant - Smart home platform. Useful for: actions.
- Assist - HA voice assistant. Useful for: the pipeline.
Architecture
Voice input (microphone)
│
▼
Whisper (STT) ──► "Turn on the light"
│
▼
Ollama (LLM) ──► Understands: light.turn_on
│
▼
Home Assistant ──► Executes: light on
│
▼
Piper (TTS) ──► "The light is now on"
│
▼
Voice output (speaker)
Setup: Completely local voice assistant
1. Docker stack
version: "3.8"
services:
whisper:
image: rhasspy/wyoming-whisper:latest
ports:
- "10300:10300"
volumes:
- whisper_data:/data
command: --model small --language de
networks:
- voice
piper:
image: rhasspy/wyoming-piper:latest
ports:
- "10200:10200"
volumes:
- piper_data:/data
command: --voice de_DE-thorsten-high
networks:
- voice
ollama:
image: ollama/ollama:latest
ports:
- "11434:11434"
volumes:
- ollama_data:/root/.ollama
networks:
- voice
homeassistant:
image: homeassistant/home-assistant:latest
ports:
- "8123:8123"
volumes:
- ha_config:/config
networks:
- voice
volumes:
whisper_data:
piper_data:
ollama_data:
ha_config:
networks:
voice:
driver: bridge
2. Configure Home Assistant
# configuration.yaml
# Wyoming integration for Whisper and Piper
# Ollama as conversation agent
ollama:
url: http://ollama:11434
In the UI:
- Settings → Devices & Services → Add Integration → Wyoming
- Whisper:
http://whisper:10300 - Piper:
http://piper:10200
- Whisper:
- Settings → Voice Assistants → Assist
- STT: Whisper
- Conversation Agent: Ollama
- TTS: Piper
3. Satellite (microphone + speaker)
For voice input, you need a satellite:
# ESP32-S3 with Micro Wake Word
# Or: Raspberry Pi with ReSpeaker
# Or: Wyoming Satellite on Linux
Practical example 1: Simple commands
# In Home Assistant: Assist commands
automation:
- alias: "Turn on light"
trigger:
- platform: conversation
command:
- "Turn on the light"
- "Light on"
- "Switch on the light"
action:
- service: light.turn_on
target:
area_id: living_room
Practical example 2: Context-based commands
automation:
- alias: "AI context command"
trigger:
- platform: conversation
command:
- "Make it cozy"
- "Cozy atmosphere"
action:
- service: light.turn_on
data:
brightness: 128
color_temp: 300
- service: media_player.play_media
data:
media_content_type: music
media_content_id: "cozy playlist"
- service: climate.set_temperature
data:
temperature: 22
Practical example 3: Answering questions
automation:
- alias: "AI question"
trigger:
- platform: conversation
command:
- "What's the weather"
- "What is the temperature"
action:
- service: ollama.generate
data:
prompt: |
Current data:
- Outside temperature: {{ states('sensor.outside_temperature') }}°C
- Weather: {{ states('weather.home') }}
- Humidity: {{ states('sensor.humidity') }}%
Answer the question: "{{ trigger.command }}"
Keep it short and friendly.
response_variable: answer
- service: tts.speak
data:
message: "{{ answer.text }}"
media_player_entity_id: media_player.kitchen
Practical example 4: Complex commands
automation:
- alias: "AI complex command"
trigger:
- platform: conversation
command:
- "Close the blinds when the sun is shining"
action:
- condition:
- condition: sun
before: sunset
after: sunrise
- service: cover.close_cover
target:
entity_id: cover.living_room
Security notes
- Everything local: No voice data leaves your network.
- Validation: Critical commands (unlock door) should require confirmation.
- Microphone access: Only authorized devices should have access.
- Logging: Record commands for debugging.
Common pitfalls
- Latency: Whisper + Ollama takes 3-7 seconds. For faster responses, use smaller models.
- German language: Whisper small is good, large is more accurate but slower.
- Misinterpretation: “Turn on the light” can affect multiple rooms. Add context.
- Microphone quality: Bad microphone equals bad transcription.
- Background noise: Loud environments interfere with recognition. Good microphone placement helps.
Further reading
- Home Assistant with AI - HA + Ollama.
- Smart Home Agent - AI agents.
- Voice Assistants - Overview.
- Whisper - STT models.
- Ollama - Model server.
- Node-RED Home Assistant - Complex workflows.
Key takeaways:
- Local voice assistant: Whisper + Ollama + Piper, completely offline.
- No data leaves your network, unlike Alexa or Google.
- Home Assistant as the action layer, Assist as the pipeline.
- For privacy-conscious smart home users.
- 3-7 seconds latency, depending on model and hardware.
FAQ
What is a local voice assistant?
Local or Alexa?
What hardware do I need?
How fast is the response?
Does German work well?
Is my voice data secure?
Does it work offline?
What if the assistant misunderstands?
Sources and further reading
- Whisper - Speech-to-text.
- Piper - Text-to-speech.
- Home Assistant Assist - Voice assistant.
- Ollama - Model server.


