Grounded AI: The Reference Architecture an FDE Deploys

Offline ingestion indexes the customer's data; the online query path grounds the model and returns cited answers

Grounded AI: The Reference Architecture an FDE Deploys Offline ingestion indexes the customer's data; the online query path grounds the model and returns cited answers Offline ingestion pipeline Online query path Customer Data · docs, tickets, DBs, wikis · Architecture component · messy + permissioned Customer Data docs, tickets, DBs, wikis messy + permissioned Ingestion · clean, chunk, embed · Offline ingestion pipeline Ingestion clean, chunk, embed Vector Store · embeddings + metadata · Offline ingestion pipeline · permission metadata Vector Store embeddings + metadata permission metadata User / Workflow · asks a question · Architecture component User / Workflow asks a question Retrieval + Rerank · top-k, permission-filtered · Online query path Retrieval + Rerank top-k, permission-filtered AI Gateway · route, cache, log, cost · Online query path AI Gateway route, cache, log, cost LLM · grounded prompt · Architecture component LLM grounded prompt Guardrails + Citations · check output, return sources · Online query path Guardrails + Citations check output, return sources permissioned sync chunks + vectors top-k query context prompt completion answer + sources Legend Database Backend Frontend Cloud External Security

Two flows

  • • Offline: index the customer's data once (and re-sync)
  • • Online: retrieve, ground, answer per request
  • • The value is the data, not the model

The hard parts the demo skipped

  • • Data is messy — cleaning and chunking drive quality
  • • Retrieval must respect permissions (a security requirement)
  • • Bad answers are usually retrieval failures, not the model

Trust by design

  • • Route via a gateway: provider-independence, cost, logging
  • • Return citations so users can verify
  • • Abstain when the answer isn't in the data