Real-Time Voice Agents: WebRTC & LiveKit
Building bidirectional, conversational voice AI systems with sub-500ms response times. Learn native audio-to-audio multimodal streaming, WebRTC transport, voice activity detection (VAD), and seamless barge-in interruption handling.
01 Cascaded Pipelines vs. Native Speech-to-Speech
ArchitectureHistorically, voice bots chained three independent components: Speech-to-Text (STT) $\rightarrow$ LLM $\rightarrow$ Text-to-Speech (TTS). Today, modern systems choose between two architectures:
Deepgram / Whisper $\rightarrow$ LLM $\rightarrow$ Cartesia / ElevenLabs. Allows hot-swapping any LLM (including custom open-weights) and granular control over tool calling, but introduces sequential network hops and latency jitter.
OpenAI Realtime API / Gemini Multimodal Live. Audio tokens directly stream in and audio tokens directly stream out over WebRTC. Captures emotional inflection, laughter, pacing, and vocal tone changes natively.
02 Building a LiveKit Voice Agent in Python
LiveKit SDKLiveKit manages media tracks, SFU routing, and client WebRTC handshakes. Here is a production voice assistant worker:
import asyncio
from livekit.agents import AutoSubscribe, JobContext, WorkerOptions, cli, llm
from livekit.agents.voice_assistant import VoiceAssistant
from livekit.plugins import openai, silero
async def entrypoint(ctx: JobContext):
# 1. Connect agent worker to the WebRTC room
await ctx.connect(auto_subscribe=AutoSubscribe.AUDIO_ONLY)
# 2. Configure Silero VAD for sub-50ms voice activity detection
vad = silero.VAD.load()
# 3. Configure low-latency voice pipeline
assistant = VoiceAssistant(
vad=vad,
stt=openai.STT(),
llm=openai.LLM(model="gpt-6-sol"),
tts=openai.TTS(voice="alloy"),
# Enable instant interruption (barge-in)
interrupt_speech_duration=0.4,
allow_interruptions=True
)
# 4. Attach domain functions for voice-driven tool calling
@assistant.fnc
def query_flight_status(flight_number: str) -> str:
"""Looks up the live departure time and gate for an airline flight."""
return f"Flight {flight_number} is on schedule, departing from Gate B22 at 3:45 PM."
# 5. Start audio processing loop and greet user
assistant.start(ctx.room)
await assistant.say("Hello! I am your autonomous airline assistant. How can I help you today?")
if __name__ == "__main__":
cli.run_app(WorkerOptions(entrypoint_fnc=entrypoint))
03 Solving the Barge-In Interruption Problem
UX EngineeringIn natural conversation, humans interrupt each other. When an agent is streaming speech and the user begins talking, the system must execute three synchronized actions within 100ms:
Immediately halt WebRTC audio track playback in the client speaker buffer so the agent stops talking instantly.
Send an abort signal to the LLM generation stream to avoid paying for tokens the user will never hear.
Record only the words the agent actually spoke before being cut off into the conversational history context window.
Frequently Asked Questions
Why is WebRTC preferred over WebSockets for voice agents? →
WebSockets run on TCP, which enforces packet retransmission and head-of-line blocking. If a packet drops over cellular networks, audio stalls. WebRTC uses UDP and RTP with adaptive jitter buffers, forward error correction (FEC), and packet loss concealment (PLC), delivering crisp real-time audio even over degraded wireless links.
Can real-time voice agents call APIs without causing awkward pauses? →
Yes, using filler speech and parallel tool execution. If an agent determines an API query will take >800ms, it streams conversational filler (e.g. "Let me check that flight for you right now...") while the asynchronous background task resolves the database payload.