• The pipeline is only as fast as its slowest stage
Everything streams
• ASR streams partial transcripts as the user speaks
• The LLM streams tokens; TTS synthesizes them as they arrive
• Streaming end-to-end is what makes the latency bearable
It's a loop
• Playback ends and capture resumes for the next turn
• VAD endpointing decides when the user has finished
• Barge-in lets the user interrupt the agent mid-speech
Data-flow diagram • Built with Archify • Create yours ↗ • Hover to trace • R route • Click to focus • +/− zoom • M radar • [/] views • P play story • T theme • E export