#Recommender Systems
Articles about Recommender Systems — exploring patterns, best practices, and real-world implementations in production systems.
8 posts tagged with recommender systems. ← All posts
Building a good model is maybe half the work; running a recommender in production is the other half. Real systems must serve in milliseconds, stay fresh as the catalog and tastes change, handle cold start gracefully, resist the filter bubbles their own optimization creates, and be monitored like any critical service. This closing post assembles everything into what it takes to run a recommender for real — and how to build one.
Building a good model is half the work; running a recommender in production is the other half. Real systems must serve in milliseconds, stay fresh as the catalog and tastes change, handle cold start gracefully, resist the filter bubbles their own optimization creates, and be monitored like any critical service. This closing post assembles everything — and gives a practical path to building one.
A recommender that scores well offline can flop in production, and a metric that looks like success can quietly harm the product. Evaluating recommenders is genuinely hard: offline metrics only approximate real behavior, the only ground truth is a live A/B test, and the very act of recommending shapes the data you learn from next. This post covers offline metrics, online testing, the gap between them, and the feedback loops that make evaluation a moving target.
A recommender that scores well offline can flop in production, and a metric that looks like success can quietly harm the product. Evaluation is genuinely hard: offline metrics only approximate real behavior, the only ground truth is a live A/B test, and the very act of recommending shapes the data you learn from next. Offline metrics, online testing, the offline-online gap, and feedback loops.
Modern large-scale recommenders — the ones running at the biggest consumer platforms — are built on neural networks. Deep learning didn't replace the core ideas (embeddings, two stages) so much as supercharge them: neural models learn richer embeddings, ingest far more features, and capture complex non-linear patterns that dot products can't. This post covers the two workhorses — two-tower retrieval and neural ranking — that power today's systems.
Modern large-scale recommenders are built on neural networks — not replacing the core ideas (embeddings, two stages) but supercharging them. This post covers the two workhorses: two-tower models for retrieval (the neural evolution of matrix factorization, built for fast ANN search) and rich neural ranking models that score the shortlist with cross-features, sequences, and multiple objectives.
You cannot run your best, most expensive model on millions of items for every request — the latency and cost are impossible. The elegant, near-universal answer is to split recommendation into two stages: a cheap, fast net that narrows millions of items to a few hundred, followed by a precise, heavier model that ranks those few. This retrieve-then-rank structure is the single most important architectural pattern in production recommenders.
You can't run your best, most expensive model on millions of items per request. The near-universal answer is to split recommendation into two stages: a cheap, fast net that narrows millions to a few hundred, then a precise, heavier model that ranks those few. This retrieve-then-rank structure is the single most important architectural pattern in production recommenders — with an interactive pipeline diagram.
The technique that defined the modern era of recommendation is deceptively simple: represent every user and every item as a short list of numbers — a vector of latent factors — such that a user's affinity for an item is just the dot product of their vectors. Matrix factorization turned recommendation into learning good embeddings, and it's the conceptual bridge from classical collaborative filtering to today's deep-learning systems.
Represent every user and item as a short vector of latent factors, and a user's affinity for an item becomes just the dot product of their vectors. Matrix factorization turned recommendation into learning good embeddings — the technique that won the Netflix Prize and the conceptual bridge from classical collaborative filtering to today's deep-learning systems and vector search.
Collaborative filtering has one crippling blind spot: it knows nothing about brand-new users or items, because they have no interactions to learn from. Content-based filtering fills that gap by recommending based on what items are rather than who interacted with them — and understanding the cold-start problem, and how each approach handles it, is key to building a recommender that works from day one.
Collaborative filtering has one crippling blind spot: it knows nothing about brand-new users or items. Content-based filtering fills that gap by recommending based on what items are rather than who interacted with them. Understanding the cold-start problem — and how each approach handles it — is key to building a recommender that works from day one, which is why most real systems are hybrids.
The most powerful idea in recommendation is also the simplest to state: you can recommend things to someone based purely on the behavior of people like them, without knowing anything about the items themselves. Collaborative filtering turns "people who liked what you liked also liked X" into an algorithm, and it's the backbone of the whole field. This post covers how it works and why it's so effective — and where it breaks.
The most powerful idea in recommendation is also the simplest: recommend things based purely on the behavior of people like you, without knowing anything about the items themselves. Collaborative filtering turns 'people who liked what you liked also liked X' into an algorithm, in user-based and item-based flavors — the backbone of the whole field, and where it breaks (cold start, sparsity).
Every time a streaming service suggests what to watch, a shop shows "you might also like," or a feed decides what you see next, a recommender system is at work. Behind that simple experience is one of the most economically important and technically rich problems in applied machine learning: from a catalog of millions, pick the handful a specific person will want, right now. This series builds recommender systems from the ground up.
Every time a service suggests what to watch, buy, or read next, a recommender system is at work — one of the most economically important and technically rich problems in applied ML: from a catalog of millions, pick the handful a specific person will want, right now. This series builds recommender systems from the ground up, starting with the problem and the two-stage architecture that structures nearly all of them.
All posts on this site are written by Pratik Dhanave, an Agentic AI Architect with 7+ years building production distributed systems, multi-agent AI platforms, and cloud-native infrastructure. About the author → Each article includes working code, architecture diagrams, and references to the specific frameworks and standards discussed. Browse all posts or explore related topics using the tag cloud above.