The Recommender Pipeline

Millions of items become a handful of recommendations: retrieve, then rank, then shape

The Recommender Pipeline Millions of items become a handful of recommendations: retrieve, then rank, then shape 01 / Sources 02 / Candidate Generation 03 / Ranking 04 / Filtering 05 / Serving Item Catalog · millions of items · 01 / Sources · corpus Item Catalog millions of items corpus User History · interactions · 01 / Sources · signals User History interactions signals Candidate Gen · retrieval (M→hundreds) · 02 / Candidate Generation · millions to hundreds Candidate Gen retrieval (M→hundreds) millions to hundreds Ranking Model · score each candidate · 03 / Ranking · precise Ranking Model score each candidate precise Filtering + Rules · dedupe, diversity, policy · 04 / Filtering · shape Filtering + Rules dedupe, diversity, policy shape Top-N Results · to the user · 05 / Serving · recommendations Top-N Results to the user recommendations item pool corpus user signals behavior shortlist shortlist ranked list ranked list top-N results Legend primary data policy / PII async batch data store

Two stages, not one

  • • You can't rank millions of items per request
  • • Candidate generation narrows the catalog to a few hundred, fast
  • • Ranking then scores that shortlist precisely

Recall then precision

  • • Retrieval optimizes for cheap, high recall (don't miss good items)
  • • Ranking optimizes for precision (order the few that matter)
  • • Different models, different objectives, chained

Shape before serving

  • • Filtering adds dedup, diversity, and business/policy rules
  • • Freshness and already-seen removal happen here
  • • The user sees a short, ordered, shaped list