AI Gateway Request Flow

Cache lookup, provider routing, transparent failover, then meter and return

AI Gateway Request Flow Cache lookup, provider routing, transparent failover, then meter and return POST /chat/completions semantic lookup miss route (within budget) timeout / 503 failover (circuit open) completion store response log + cost + trace 200 completion Cache check Route + failover Meter + return Application · agent / backend · Sequence participant Application agent / backend AI Gateway · one control point · Sequence participant AI Gateway one control point Cache · exact + semantic · Sequence participant Cache exact + semantic Primary · e.g. OpenAI · Sequence participant Primary e.g. OpenAI Fallback · e.g. Anthropic · Sequence participant Fallback e.g. Anthropic Observability · logs · cost · traces · Sequence participant Observability logs · cost · traces Legend request return security async trace

Cache first

  • • Semantic lookup runs before any provider call
  • • A hit returns instantly, skipping the model entirely
  • • Cuts both cost and latency on repeat queries

Transparent failover

  • • A primary timeout/error trips the circuit breaker
  • • The same request retries on a fallback provider
  • • The application sees one call, never the outage

Metered and returned

  • • The response is cached for next time
  • • The call is logged with cost and a trace
  • • One clean response returns to the caller