Recommender Systems from the Ground Up

Recommender systems from the ground up — the recommendation problem (match people to items from a huge catalog, personally and in real time, from sparse noisy interaction data; explicit vs implicit feedback; scale, personalization, cold start, latency), collaborative filtering (recommend from the behavior of similar users/items using only the interaction matrix; user-based vs item-based; similarity measures; serendipity, sparsity, and cold-start limits), content-based filtering and cold start (recommend by item features; handles new items and niche users; the three faces of cold start and the toolkit — content features, onboarding, popular fallbacks, exploration; why hybrids win), matrix factorization and embeddings (latent factors, the dot-product model, the Netflix Prize; users and items as embeddings in a shared space — the seed of the modern field), the two-stage architecture (retrieve-then-rank: candidate generation narrows millions to hundreds via embedding ANN search, then a precise ranker orders them, then filtering/shaping — the same pattern as search and RAG), deep-learning recommenders (two-tower retrieval and rich neural ranking with cross-features, sequences, and multi-objective), evaluating recommenders (ranking metrics like NDCG/Precision@K, online A/B testing as ground truth, the offline-online gap and exposure bias, feedback loops), and production recommenders (millisecond serving, freshness, cold start and the long tail, filter bubbles and responsible recommendation, and a practical build path). Includes an interactive archify pipeline diagram. Grounded in Wikipedia, Google's ML recommendation course, and the Netflix Prize.

8 parts · written by Pratik Dhanave. Start with Part 1 →

← All series · All posts

Part 1 · ·6 min read

The Recommendation Problem

Every time a streaming service suggests what to watch, a shop shows "you might also like," or a feed decides what you see next, a recommender system is at work. Behind that simple experience is one of the most economically important and technically rich problems in applied machine learning: from a catalog of millions, pick the handful a specific person will want, right now. This series builds recommender systems from the ground up.

Every time a service suggests what to watch, buy, or read next, a recommender system is at work — one of the most economically important and technically rich problems in applied ML: from a catalog of millions, pick the handful a specific person will want, right now. This series builds recommender systems from the ground up, starting with the problem and the two-stage architecture that structures nearly all of them.

Part 2 · ·5 min read

Collaborative Filtering

The most powerful idea in recommendation is also the simplest to state: you can recommend things to someone based purely on the behavior of people like them, without knowing anything about the items themselves. Collaborative filtering turns "people who liked what you liked also liked X" into an algorithm, and it's the backbone of the whole field. This post covers how it works and why it's so effective — and where it breaks.

The most powerful idea in recommendation is also the simplest: recommend things based purely on the behavior of people like you, without knowing anything about the items themselves. Collaborative filtering turns 'people who liked what you liked also liked X' into an algorithm, in user-based and item-based flavors — the backbone of the whole field, and where it breaks (cold start, sparsity).

Part 3 · ·6 min read

Content-Based Filtering and the Cold-Start Problem

Collaborative filtering has one crippling blind spot: it knows nothing about brand-new users or items, because they have no interactions to learn from. Content-based filtering fills that gap by recommending based on what items are rather than who interacted with them — and understanding the cold-start problem, and how each approach handles it, is key to building a recommender that works from day one.

Collaborative filtering has one crippling blind spot: it knows nothing about brand-new users or items. Content-based filtering fills that gap by recommending based on what items are rather than who interacted with them. Understanding the cold-start problem — and how each approach handles it — is key to building a recommender that works from day one, which is why most real systems are hybrids.

Part 4 · ·6 min read

Matrix Factorization and Embeddings

The technique that defined the modern era of recommendation is deceptively simple: represent every user and every item as a short list of numbers — a vector of latent factors — such that a user's affinity for an item is just the dot product of their vectors. Matrix factorization turned recommendation into learning good embeddings, and it's the conceptual bridge from classical collaborative filtering to today's deep-learning systems.

Represent every user and item as a short vector of latent factors, and a user's affinity for an item becomes just the dot product of their vectors. Matrix factorization turned recommendation into learning good embeddings — the technique that won the Netflix Prize and the conceptual bridge from classical collaborative filtering to today's deep-learning systems and vector search.

Part 5 · ·6 min read

The Two-Stage Architecture

You cannot run your best, most expensive model on millions of items for every request — the latency and cost are impossible. The elegant, near-universal answer is to split recommendation into two stages: a cheap, fast net that narrows millions of items to a few hundred, followed by a precise, heavier model that ranks those few. This retrieve-then-rank structure is the single most important architectural pattern in production recommenders.

You can't run your best, most expensive model on millions of items per request. The near-universal answer is to split recommendation into two stages: a cheap, fast net that narrows millions to a few hundred, then a precise, heavier model that ranks those few. This retrieve-then-rank structure is the single most important architectural pattern in production recommenders — with an interactive pipeline diagram.

Part 6 · ·6 min read

Deep Learning Recommenders

Modern large-scale recommenders — the ones running at the biggest consumer platforms — are built on neural networks. Deep learning didn't replace the core ideas (embeddings, two stages) so much as supercharge them: neural models learn richer embeddings, ingest far more features, and capture complex non-linear patterns that dot products can't. This post covers the two workhorses — two-tower retrieval and neural ranking — that power today's systems.

Modern large-scale recommenders are built on neural networks — not replacing the core ideas (embeddings, two stages) but supercharging them. This post covers the two workhorses: two-tower models for retrieval (the neural evolution of matrix factorization, built for fast ANN search) and rich neural ranking models that score the shortlist with cross-features, sequences, and multiple objectives.

Part 7 · ·6 min read

Evaluating Recommenders

A recommender that scores well offline can flop in production, and a metric that looks like success can quietly harm the product. Evaluating recommenders is genuinely hard: offline metrics only approximate real behavior, the only ground truth is a live A/B test, and the very act of recommending shapes the data you learn from next. This post covers offline metrics, online testing, the gap between them, and the feedback loops that make evaluation a moving target.

A recommender that scores well offline can flop in production, and a metric that looks like success can quietly harm the product. Evaluation is genuinely hard: offline metrics only approximate real behavior, the only ground truth is a live A/B test, and the very act of recommending shapes the data you learn from next. Offline metrics, online testing, the offline-online gap, and feedback loops.

Part 8 · ·7 min read

Production Recommenders

Building a good model is maybe half the work; running a recommender in production is the other half. Real systems must serve in milliseconds, stay fresh as the catalog and tastes change, handle cold start gracefully, resist the filter bubbles their own optimization creates, and be monitored like any critical service. This closing post assembles everything into what it takes to run a recommender for real — and how to build one.

Building a good model is half the work; running a recommender in production is the other half. Real systems must serve in milliseconds, stay fresh as the catalog and tastes change, handle cold start gracefully, resist the filter bubbles their own optimization creates, and be monitored like any critical service. This closing post assembles everything — and gives a practical path to building one.

This series is part of a larger body of work by Pratik Dhanave, an Agentic AI Architect writing about production AI systems, distributed systems, and cloud-native engineering. Explore all course series, browse every post, or find topics via the tag index.