Evaluating Recommenders
A recommender that scores well offline can flop in production, and a metric that looks like success can quietly harm the product. Evaluating recommenders is genuinely hard: offline metrics only approximate real behavior, the only ground truth is a live A/B test, and the very act of recommending shapes the data you learn from next. This post covers offline metrics, online testing, the gap between them, and the feedback loops that make evaluation a moving target.
A recommender that scores well offline can flop in production, and a metric that looks like success can quietly harm the product. Evaluation is genuinely hard: offline metrics only approximate real behavior, the only ground truth is a live A/B test, and the very act of recommending shapes the data you learn from next. Offline metrics, online testing, the offline-online gap, and feedback loops.