AI Gateway Reference Architecture

One control point between your applications and every model provider

AI Gateway Reference Architecture One control point between your applications and every model provider AI Gateway Applications · agents · chatbots · backends · Architecture component Applications agents · chatbots · backends Unified API · one OpenAI-style interface · AI Gateway Unified API one OpenAI-style interface Router · model routing + load balancing · AI Gateway Router model routing + load balancing Cache · exact + semantic · AI Gateway Cache exact + semantic Rate Limiter · quotas + budgets · AI Gateway Rate Limiter quotas + budgets Observability · logs · traces · cost · AI Gateway Observability logs · traces · cost Fallback / Retry · circuit breaker · AI Gateway Fallback / Retry circuit breaker OpenAI · GPT models · Architecture component OpenAI GPT models Anthropic · Claude models · Architecture component Anthropic Claude models Google Vertex · Gemini models · Architecture component Google Vertex Gemini models Self-hosted · vLLM / open weights · Architecture component Self-hosted vLLM / open weights one request check enforce emit on failure Legend External Backend Database Security

One interface

  • • Apps call a single OpenAI-style API
  • • Gateway translates to each provider's format
  • • Swap models without changing app code

Reliability + cost

  • • Route + load-balance across providers/keys
  • • Fallback + circuit breaker on outages
  • • Exact + semantic caching cuts cost/latency

Governance

  • • Rate limits, quotas, and spend budgets
  • • Central logging, tracing, and cost metering
  • • One choke point for policy and guardrails