aiagent.org logo aiagent.org
🎙️ Sub-500ms Audio Engineering

Real-Time Voice Agents: WebRTC & LiveKit

Building bidirectional, conversational voice AI systems with sub-500ms response times. Learn native audio-to-audio multimodal streaming, WebRTC transport, voice activity detection (VAD), and seamless barge-in interruption handling.

Latency: Sub-500ms Glass-to-Glass
Protocols: WebRTC vs WebSockets
Engines: LiveKit, OpenAI Realtime, Gemini Live

01 Cascaded Pipelines vs. Native Speech-to-Speech

Architecture

Historically, voice bots chained three independent components: Speech-to-Text (STT) $\rightarrow$ LLM $\rightarrow$ Text-to-Speech (TTS). Today, modern systems choose between two architectures:

Cascaded Pipeline ~1200ms Latency

Deepgram / Whisper $\rightarrow$ LLM $\rightarrow$ Cartesia / ElevenLabs. Allows hot-swapping any LLM (including custom open-weights) and granular control over tool calling, but introduces sequential network hops and latency jitter.

Native Speech-to-Speech <450ms Latency

OpenAI Realtime API / Gemini Multimodal Live. Audio tokens directly stream in and audio tokens directly stream out over WebRTC. Captures emotional inflection, laughter, pacing, and vocal tone changes natively.

02 Building a LiveKit Voice Agent in Python

LiveKit SDK

LiveKit manages media tracks, SFU routing, and client WebRTC handshakes. Here is a production voice assistant worker:

# Install: pip install livekit-agents livekit-plugins-openai livekit-plugins-silero
import asyncio
from livekit.agents import AutoSubscribe, JobContext, WorkerOptions, cli, llm
from livekit.agents.voice_assistant import VoiceAssistant
from livekit.plugins import openai, silero

async def entrypoint(ctx: JobContext):
    # 1. Connect agent worker to the WebRTC room
    await ctx.connect(auto_subscribe=AutoSubscribe.AUDIO_ONLY)

    # 2. Configure Silero VAD for sub-50ms voice activity detection
    vad = silero.VAD.load()

    # 3. Configure low-latency voice pipeline
    assistant = VoiceAssistant(
        vad=vad,
        stt=openai.STT(),
        llm=openai.LLM(model="gpt-6-sol"),
        tts=openai.TTS(voice="alloy"),
        # Enable instant interruption (barge-in)
        interrupt_speech_duration=0.4,
        allow_interruptions=True
    )

    # 4. Attach domain functions for voice-driven tool calling
    @assistant.fnc
    def query_flight_status(flight_number: str) -> str:
        """Looks up the live departure time and gate for an airline flight."""
        return f"Flight {flight_number} is on schedule, departing from Gate B22 at 3:45 PM."

    # 5. Start audio processing loop and greet user
    assistant.start(ctx.room)
    await assistant.say("Hello! I am your autonomous airline assistant. How can I help you today?")

if __name__ == "__main__":
    cli.run_app(WorkerOptions(entrypoint_fnc=entrypoint))

03 Solving the Barge-In Interruption Problem

UX Engineering

In natural conversation, humans interrupt each other. When an agent is streaming speech and the user begins talking, the system must execute three synchronized actions within 100ms:

1. Audio Truncation

Immediately halt WebRTC audio track playback in the client speaker buffer so the agent stops talking instantly.

2. Cancel Token Stream

Send an abort signal to the LLM generation stream to avoid paying for tokens the user will never hear.

3. Transcript Slicing

Record only the words the agent actually spoke before being cut off into the conversational history context window.

Frequently Asked Questions

Why is WebRTC preferred over WebSockets for voice agents? →

WebSockets run on TCP, which enforces packet retransmission and head-of-line blocking. If a packet drops over cellular networks, audio stalls. WebRTC uses UDP and RTP with adaptive jitter buffers, forward error correction (FEC), and packet loss concealment (PLC), delivering crisp real-time audio even over degraded wireless links.

Can real-time voice agents call APIs without causing awkward pauses? →

Yes, using filler speech and parallel tool execution. If an agent determines an API query will take >800ms, it streams conversational filler (e.g. "Let me check that flight for you right now...") while the asynchronous background task resolves the database payload.