The Voice Agent Pipeline

Audio becomes text becomes a response becomes audio — a real-time cascade

The Voice Agent Pipeline Audio becomes text becomes a response becomes audio — a real-time cascade 01 / Capture 02 / Transcribe 03 / Reason 04 / Synthesize 05 / Play Microphone · audio stream · 01 / Capture · 16kHz PCM Microphone audio stream 16kHz PCM VAD · endpointing · 01 / Capture · when done? VAD endpointing when done? ASR / STT · Whisper-style · 02 / Transcribe · streaming ASR / STT Whisper-style streaming LLM · the brain · 03 / Reason · streaming LLM the brain streaming TTS · voice synthesis · 04 / Synthesize · streaming TTS voice synthesis streaming Speaker · audio out · 05 / Play · playback Speaker audio out playback audio frames raw audio speech segment endpointed audio transcript text response tokens text audio chunks synthesized audio Legend primary data policy / PII async batch data store

A cascade of transforms

  • • Audio → text (ASR) → response (LLM) → audio (TTS)
  • • Each stage is a separate model or service
  • • The pipeline is only as fast as its slowest stage

Everything streams

  • • ASR streams partial transcripts as the user speaks
  • • The LLM streams tokens; TTS synthesizes them as they arrive
  • • Streaming end-to-end is what makes the latency bearable

It's a loop

  • • Playback ends and capture resumes for the next turn
  • • VAD endpointing decides when the user has finished
  • • Barge-in lets the user interrupt the agent mid-speech