Skip to content
VDAI with VD

May 13, 2026 · 8 min read

Building a Local-First Voice AI Assistant with LiveKit Agents

How I built a real-time, hybrid voice AI: VAD, STT, and LLM run locally, while default Edge-TTS sends response text to Microsoft Edge's online TTS service.

  • LiveKit
  • Voice AI
  • FastAPI
  • Whisper
  • Ollama
  • WebRTC
  • Next.js
  • Local-First
  • Edge-TTS
  • Real-Time Audio

Voice assistants are everywhere — Siri, Alexa, Google Assistant. But they all share the same fundamental trade-off: your audio leaves your device, gets processed on someone else's servers, and your conversation history lives on infrastructure you don't control.

I wanted a system that keeps the compute-heavy speech and reasoning path under my control while retaining the freedom to swap any component. So I built Voice AI Demo — an open-source, local-first voice assistant powered by LiveKit Agents. It is a hybrid deployment: VAD, STT, and LLM run locally, while the default TTS provider uses an online synthesis service.

In this post I'll walk through the architecture, the four-stage voice pipeline, how LiveKit orchestrates WebRTC media, and the production patterns that make it reliable.

The Problem

Building a local-first voice AI assistant means solving five hard problems:

  1. Real-time audio transport — You can't use HTTP for streaming audio. You need WebRTC with low-latency media channels
  2. Voice Activity Detection — When does the user start speaking? When do they stop? Naive approaches break on background noise
  3. Speech-to-Text — Transcribing audio locally requires a model and runtime that fit the interaction loop
  4. Language Understanding — The LLM needs to produce useful responses while keeping the compute and data boundaries explicit
  5. Text-to-Speech — Synthesizing natural-sounding audio while making the provider's network boundary explicit

LiveKit Agents solves problems 1 and 5 elegantly. The rest is about choosing the right models and wiring them together resiliently.

Architecture Overview

Voice AI Demo Architecture

At a glance:

Browser → LiveKit (WebRTC) → Agent Worker → VAD → STT → LLM → TTS
                  ↕                         ↕
            LiveKit Server              Ollama (local)

Core components:

LayerTechnologyRole
FrontendNext.js 15, React 19, Agents UIVoice interface with mic controls, audio visualizer, chat transcript
Agent ServerPython 3.12, FastAPI, LiveKit Agents SDKJWT token generation, health check, pipeline orchestration
MediaLiveKit Server (Docker/Go)WebRTC SFU, room management, job dispatch
VADSilero VADSpeech segment and utterance boundary detection
STTfaster-whisper (large-v3-turbo)Local speech-to-text transcription
LLMOllama (llama3.2:3b)Local language model via OpenAI-compatible API
TTSEdge-TTSOnline Microsoft Edge text-to-speech service (default, swappable)

How LiveKit Agents Works

LiveKit Agents is the backbone of this system. It provides:

  • WebRTC SFU — A Selective Forwarding Unit that routes media between browsers and agent workers
  • PipelineAgent — A high-level abstraction that processes audio through config stages (VAD → STT → LLM → TTS)
  • Job dispatch — When a user connects, LiveKit automatically dispatches a job to an available agent worker
  • JWT authentication — Secure token-based room access with configurable permissions

The data flow is elegant:

  1. User opens http://localhost:3000 in a browser
  2. Frontend calls GET /token on FastAPI → receives a signed LiveKit JWT
  3. Browser connects to LiveKit Server via WebRTC using the JWT
  4. LiveKit dispatches a job to the Agent Worker (background thread)
  5. Agent processes audio through the voice pipeline
  6. Response audio streams back through LiveKit → browser plays it in real-time

The 4-Stage Voice Pipeline

Stage 1: VAD (Voice Activity Detection)

Silero VAD detects when speech starts and ends. This is critical for two reasons:

  • Utterance detection — Know when the user has finished speaking so the pipeline can start processing
  • Noise filtering — Silence and background noise are discarded before they reach the STT model

Silero is pre-trained and small (~1.7MB), and this build runs it on CPU. It operates on 30ms audio frames and returns a probability score between 0 and 1 for each frame.

Stage 2: STT (Speech-to-Text)

faster-whisper is a CTranslate2 implementation of OpenAI's Whisper model. This build uses it for local speech-to-text within the LiveKit pipeline.

The pipeline uses large-v3-turbo and enables the available Metal acceleration path on Apple Silicon. Audio is passed to the local model, and the resulting transcript is returned to the agent. This article does not claim a comparative speed or accuracy benchmark for that configuration.

The first run downloads the model (~3GB). Subsequent runs reuse the local cache.

Stage 3: LLM (Language Model)

Ollama serves Llama 3.2 3B locally via an OpenAI-compatible API. The agent sends transcribed text along with a system prompt, and Ollama streams tokens back.

The implementation consumes Ollama's streaming response through the OpenAI-compatible interface, so the configured model can be replaced without changing the rest of the pipeline. Runtime behavior varies with hardware, model, and configuration; this article does not assert a portable throughput or typical response-time benchmark.

The magic is that you can swap to any OpenAI-compatible provider at runtime. Just change .env:

LLM_PROVIDER=ollama  # or openai_compatible
LLM_MODEL=llama3.2:3b  # or llama-3.3-70b-versatile (Groq)
LLM_BASE_URL=http://localhost:11434/v1  # or https://api.groq.com/openai/v1

Stage 4: TTS (Text-to-Speech)

The default Edge-TTS provider sends response text to Microsoft Edge's online text-to-speech service. It produces natural-sounding speech with multiple voice options (en-US-AriaNeural is the default), but synthesis does not happen on the host.

The synthesized audio streams back into LiveKit's audio track for browser playback. The TTS provider remains swappable through configuration, preserving the same pipeline interface when a deployment selects a different service. Sub-2s voice-to-voice latency was observed in local development on Apple Silicon; it is not a controlled cross-hardware benchmark and not a guarantee.

Safe Wrappers: Error Resilience

Each pipeline stage has a Safe wrapper that catches failures and provides graceful fallbacks:

class SafeSTT:
    async def transcribe(self, audio: AudioFrame) -> str:
        try:
            return await self.stt.transcribe(audio)
        except Exception as e:
            logger.error(f"STT failed: {e}")
            return ""  # Empty transcript — LLM handles gracefully

This means a single model failure doesn't crash the entire conversation. If Whisper fails, the user gets a "I didn't catch that" response. If Ollama fails, TTS gets a fallback message. The conversation keeps going.

Frontend: Agents UI

The Next.js frontend uses LiveKit's Agents UI — a shadcn-based React component library:

  • Audio visualizer — Real-time waveform display of microphone input
  • Chat transcript — Scrollable conversation history with timestamps
  • Mic controls — Mute/unmute, push-to-talk, and connection status
  • Dark theme — Optimized for the voice-first interface

The frontend is deliberately minimal. The complexity lives in the agent pipeline, not the UI.

Provider Switching

The most powerful feature: every component of the pipeline is swappable at runtime. Here are three example configurations:

Local-first hybrid (default):

STT_PROVIDER=whisper
LLM_PROVIDER=ollama
TTS_PROVIDER=edge-tts

VAD, STT, and LLM run locally. Default Edge-TTS sends the generated response text to Microsoft Edge's online TTS service, so this configuration requires network access even though Edge-TTS does not require an API key.

Hosted STT alternative:

STT_PROVIDER=openai_compatible
STT_MODEL=whisper-1
STT_BASE_URL=https://api.openai.com/v1
# Add OPENAI_API_KEY to .env

This option routes transcription through OpenAI's Whisper API. It requires an API key and moves audio across an external network boundary; runtime behavior depends on the network and workload.

Hosted LLM hybrid:

LLM_PROVIDER=openai_compatible
LLM_MODEL=llama-3.3-70b-versatile
LLM_BASE_URL=https://api.groq.com/v1
# Add GROQ_API_KEY to .env

This variant sends LLM prompts to Groq, keeps Whisper STT on the host, and continues to use Microsoft Edge's online service for speech synthesis.

Quick Start

Start each part of the local stack in this order:

# 1. Start LiveKit Server (Docker)
cp .env.example .env
docker compose up -d

# 2. Pull LLM model
ollama pull llama3.2:3b

# 3. Start Agent Service
cd agent
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py

# 4. Start Frontend
cd frontend
npm install
npm run dev

Open http://localhost:3000 and click "Start audio". You're talking to a local-first voice AI with an online TTS boundary.

Key Takeaways

  1. LiveKit Agents is the right abstraction — It handles WebRTC complexity, media routing, job dispatch, and pipeline orchestration so you focus on the AI models
  2. Local-first pipelines can support real-time interaction — Local Whisper and Llama inference, paired with default online Edge-TTS, showed sub-2s voice-to-voice latency in local development on Apple Silicon; this is not a controlled cross-hardware benchmark or guarantee
  3. Provider swappability is essential — The ability to switch STT, LLM, or TTS at runtime (not rebuild time) makes the system adaptable to any environment
  4. Safe wrappers prevent cascade failures — A single model crash should never take down the conversation. Graceful degradation keeps the UX intact
  5. Docker for the hard parts — LiveKit Server runs in Docker. The agent and frontend run natively. This separation keeps service boundaries explicit and deployment flexible

Code & Resources

  • Full source code: github.com/aiwithvd/voiceai
  • Docker setup: One-command LiveKit server with docker compose up -d
  • Testing: 15 agent tests + 4 frontend tests + --test-mode for CI
  • Models: Silero VAD, faster-whisper, and Ollama run locally; the default Edge-TTS provider uses Microsoft Edge's online service

Want to build production voice AI systems? Let's connect. I advise teams on real-time AI architecture, voice pipeline design, and self-hosted deployment strategies.

Written by

Vishvdeep Dashadiya

Lead AI Engineer. Agentic AI, real-time ML systems, and cloud-native infrastructure.

About me

Keep reading

More posts

September 3, 2026 · 7 min read

It's Not Adaptive RAG: Why I Let the Human Choose the Route

I built a RAG service for a regulated finance and legal practice and called it adaptive-rag. It isn't. There is no LLM router picking a path. Vector search always runs, web search is added only when a person flips a toggle, and four self-reflection gates with hard-capped loops decide whether the answer is good enough to ship. Here is why that design beat the classic one.

RAGLangGraph

April 13, 2026 · 14 min read

OpenClaw: A Self-Hosted AI Assistant with Ollama, Telegram & Discord

Install OpenClaw, wire Anthropic/OpenAI/Google or a local Ollama model, control it from Telegram and Discord, extend it with skills.sh, and turn it into a business gateway for lead qualification, customer support, and agent-ecosystem management.

OpenClawOllama

Working on something like this?

Tell me about it. I reply within one working day with a first take and no sales pitch.