The Two-Stage Architecture
You cannot run your best, most expensive model on millions of items for every request — the latency and cost are impossible. The elegant, near-universal answer is to split recommendation into two stages: a cheap, fast net that narrows millions of items to a few hundred, followed by a precise, heavier model that ranks those few. This retrieve-then-rank structure is the single most important architectural pattern in production recommenders.
You can't run your best, most expensive model on millions of items per request. The near-universal answer is to split recommendation into two stages: a cheap, fast net that narrows millions to a few hundred, then a precise, heavier model that ranks those few. This retrieve-then-rank structure is the single most important architectural pattern in production recommenders — with an interactive pipeline diagram.