Skip to content
BotServBotServ
Voice AssistantLocalWhisperOllamaPiperSmart Home

Local Voice Assistant for Smart Home

Local voice assistant for smart home. Whisper + Ollama + Piper, completely offline, no cloud.

S

schutzgeist

5 min read
Local Voice Assistant for Smart Home

Local Voice Assistant for Smart Home

What this article covers

  • How to build a completely local voice assistant.
  • Whisper (STT) + Ollama (comprehension) + Piper (TTS).
  • Connecting the assistant to Home Assistant.
  • Practical examples for lights, climate, music, and security.
  • Best practices for latency, accuracy, and privacy.

Introduction: Local voice assistants explained

A local voice assistant understands voice commands without the cloud. You speak, Whisper transcribes, Ollama understands and decides, Piper responds. Your data never leaves your network, unlike Alexa or Google Home.

This article is for users who want to build a local voice assistant. For fundamentals, check out Home Assistant with AI and Voice Assistants.

Why do you need a local voice assistant?

Alexa and Google Home send your voice commands to the cloud. A local assistant processes everything on your server. “Turn on the light” gets transcribed, understood, and executed locally. No data leaves, no ads, full control.

How a local voice assistant works

Microphone → Whisper (STT) → Ollama (LLM) → Home Assistant (action) → Piper (TTS) → Speaker. All local, no cloud.

The core idea: voice control without data transmission.

Who is this article for?

  • Privacy-conscious users who don’t want cloud voice assistants.
  • Smart home tinkerers building local voice control.
  • Home Assistant users running Assist locally.
  • Developers building custom voice assistants.

Key terms

  • Whisper - OpenAI’s speech-to-text (local). Useful for: voice input.
  • Ollama - Local model server. Useful for: comprehension.
  • Piper - Text-to-speech (local). Useful for: voice output.
  • Wyoming - Protocol for voice services. Useful for: integration.
  • Home Assistant - Smart home platform. Useful for: actions.
  • Assist - HA voice assistant. Useful for: the pipeline.

Architecture

Voice input (microphone)
    │
    ▼
Whisper (STT) ──► "Turn on the light"
    │
    ▼
Ollama (LLM) ──► Understands: light.turn_on
    │
    ▼
Home Assistant ──► Executes: light on
    │
    ▼
Piper (TTS) ──► "The light is now on"
    │
    ▼
Voice output (speaker)

Setup: Completely local voice assistant

1. Docker stack

version: "3.8"

services:
  whisper:
    image: rhasspy/wyoming-whisper:latest
    ports:
      - "10300:10300"
    volumes:
      - whisper_data:/data
    command: --model small --language de
    networks:
      - voice

  piper:
    image: rhasspy/wyoming-piper:latest
    ports:
      - "10200:10200"
    volumes:
      - piper_data:/data
    command: --voice de_DE-thorsten-high
    networks:
      - voice

  ollama:
    image: ollama/ollama:latest
    ports:
      - "11434:11434"
    volumes:
      - ollama_data:/root/.ollama
    networks:
      - voice

  homeassistant:
    image: homeassistant/home-assistant:latest
    ports:
      - "8123:8123"
    volumes:
      - ha_config:/config
    networks:
      - voice

volumes:
  whisper_data:
  piper_data:
  ollama_data:
  ha_config:

networks:
  voice:
    driver: bridge

2. Configure Home Assistant

# configuration.yaml
# Wyoming integration for Whisper and Piper

# Ollama as conversation agent
ollama:
  url: http://ollama:11434

In the UI:

  1. Settings → Devices & Services → Add Integration → Wyoming
    • Whisper: http://whisper:10300
    • Piper: http://piper:10200
  2. Settings → Voice Assistants → Assist
    • STT: Whisper
    • Conversation Agent: Ollama
    • TTS: Piper

3. Satellite (microphone + speaker)

For voice input, you need a satellite:

# ESP32-S3 with Micro Wake Word
# Or: Raspberry Pi with ReSpeaker
# Or: Wyoming Satellite on Linux

Practical example 1: Simple commands

# In Home Assistant: Assist commands
automation:
  - alias: "Turn on light"
    trigger:
      - platform: conversation
        command:
          - "Turn on the light"
          - "Light on"
          - "Switch on the light"
    action:
      - service: light.turn_on
        target:
          area_id: living_room

Practical example 2: Context-based commands

automation:
  - alias: "AI context command"
    trigger:
      - platform: conversation
        command:
          - "Make it cozy"
          - "Cozy atmosphere"
    action:
      - service: light.turn_on
        data:
          brightness: 128
          color_temp: 300
      - service: media_player.play_media
        data:
          media_content_type: music
          media_content_id: "cozy playlist"
      - service: climate.set_temperature
        data:
          temperature: 22

Practical example 3: Answering questions

automation:
  - alias: "AI question"
    trigger:
      - platform: conversation
        command:
          - "What's the weather"
          - "What is the temperature"
    action:
      - service: ollama.generate
        data:
          prompt: |
            Current data:
            - Outside temperature: {{ states('sensor.outside_temperature') }}°C
            - Weather: {{ states('weather.home') }}
            - Humidity: {{ states('sensor.humidity') }}%

            Answer the question: "{{ trigger.command }}"
            Keep it short and friendly.
        response_variable: answer
      - service: tts.speak
        data:
          message: "{{ answer.text }}"
          media_player_entity_id: media_player.kitchen

Practical example 4: Complex commands

automation:
  - alias: "AI complex command"
    trigger:
      - platform: conversation
        command:
          - "Close the blinds when the sun is shining"
    action:
      - condition:
          - condition: sun
            before: sunset
            after: sunrise
      - service: cover.close_cover
        target:
          entity_id: cover.living_room

Security notes

  • Everything local: No voice data leaves your network.
  • Validation: Critical commands (unlock door) should require confirmation.
  • Microphone access: Only authorized devices should have access.
  • Logging: Record commands for debugging.

Common pitfalls

  • Latency: Whisper + Ollama takes 3-7 seconds. For faster responses, use smaller models.
  • German language: Whisper small is good, large is more accurate but slower.
  • Misinterpretation: “Turn on the light” can affect multiple rooms. Add context.
  • Microphone quality: Bad microphone equals bad transcription.
  • Background noise: Loud environments interfere with recognition. Good microphone placement helps.

Further reading

Key takeaways:

  • Local voice assistant: Whisper + Ollama + Piper, completely offline.
  • No data leaves your network, unlike Alexa or Google.
  • Home Assistant as the action layer, Assist as the pipeline.
  • For privacy-conscious smart home users.
  • 3-7 seconds latency, depending on model and hardware.

FAQ

What is a local voice assistant?

A voice assistant that runs completely on your server: Whisper for STT, Ollama for comprehension, Piper for TTS. No cloud, no data transmission.

Local or Alexa?

Local for privacy and control. Alexa for convenience and music. Local is more complex but completely private.

What hardware do I need?

Server: 8-16 GB VRAM for Whisper + Ollama. Satellite: ESP32-S3 or Raspberry Pi for microphone + speaker.

How fast is the response?

3-7 seconds from speech to action. Whisper small: 1-2s, Ollama 8B: 2-5s, Piper: <1s. For faster responses, use smaller models.

Does German work well?

Yes, Whisper small and large support German. Piper with German voice (de_DE-thorsten-high). Ollama with German prompts.

Is my voice data secure?

Yes, completely. All processing happens locally. No voice data leaves your network. This is the main advantage over cloud assistants.

Does it work offline?

Yes, completely. No internet connection needed. Whisper, Ollama, and Piper run locally. You only need internet for updates.

What if the assistant misunderstands?

AI can misinterpret commands. For critical actions (unlock door), add confirmation. Precise commands and context help.

Sources and further reading

Back to Blog
Share:

Related Posts