Welcome back to Mining of Massive Datasets. In the previous lecture we looked at collaborative filtering approaches for recommender systems. In this section, we're going to look at how to evaluate a recommender system, to make sure that it's doing a good job. Let's look at an example. Suppose we have a set of users and movies. Users on the y axis, movies on the x axis. And here's a utility matrix that gives you know, so, where some of the rating values are known and some are unknown. And ratings are a scale from one through five. Now the common evaluation methodology is to take a piece of the matrix and treat that as a test set. Now we know these ratings, but there going to withhold them from the algorithm that they are going to use. So we'll call this the test data set, or the withheld ratings. And we'll for the purposes of the of the algorithm these are going to treated as the same as unknown ratings, or the blank ratings. But in fact we know what these ratings are, so we can use algorithms to predict these ratings. And then compare them against the actual ratings and see good the algorithm performed. So the the trick is to compare the predictions against the withheld known ratings of the test set T. And the most common measure is is a measure called the root-mean-square error, or RMSE for short. And it's very simply defined. You know as, you know, as the some of the squares of the deviations from the actual ratings and the predicted ratings. So in this case Rxi star is the actual rating for user X-N item i. Rxi is the rating that's predicted by, by our algorithm. We're going to take the difference between those two and square it, and sum of those squares across all the withheld ratings, and divided by the total number of withheld ratings which is N, and just take the square root of that. So, this is the root mean square error, or RMSE and it's the most commonly used measure to evaluate a, a collaborative filtering system. Now, while algorithm is a very very simple measure, it does have a few, few problems. And, one of the, one of the most common problems, is that there's this narrow focus on accuracy, or RMSE sometimes misses the point of why we implement recommend, recommend a system in the first place. But the goal of the recommended system is not really to predict the the rating of a user for a given item, but to recommend items to a user that the, the user might buy or might view or might listen to. And therefore sometimes when you predict just the, just the items with the highest Scores are the highest similarity you might end up with problems. The first problem we run into, is the problem of prediction diversity, which means that all the predictions are too similar to each other. For example, let's say the user liked the first Harry Potter movie. Now the set of most similar items or the set of items with the best scores might be the, the other Harry Potter movies. So all the predictions might therefore be Harry Potter movies. Which is you know rather a non diverse set of recommendations and might not introduce the user to any new and interesting items. The second problem is one of context. Now, the user might be the same user might be operating in different contexts and might want different items in those, in those contexts. For example, let's say a user is travelling to to Patagonia. He might end up shopping for lots of travel books on Patagonia. Then once he, once he's back from Patagonia the s the set of is similar, the set of recommended items will include bo you know more books on Patagonia. But the user is not interested in any more, because the user is in a very different context. Finally, the order of predictions also turns out to be hugely important. For example, when there's a series of books or movies you want to predict, you want to recommend items that are earlier in the sequence are ahead of items that are later in the sequence. Note also that in practice, we only care to predict high ratings. Right? So, we've sort of set up this whole problem of, predicting the, rating of a user X for an item I. But in practice we don't really care what the rating of user x for item i is. And that rating value is low and the user really doesn't like the item. We are only going to recommend items to the user that the user really cares about. And so we only care to predict high ratings, not to predict low ratings. And the RMSE that we've sort of, defined, routine square error, actually doesn't distinguish between high ratings and low ratings. It only looks at the difference between the actual ratings and the predicted ratings. So, in fact we might come up with a method that works well in practice that recommends very good items to users. But actually mix bad errors on items that the user doesn't really like. Now this algorithm might work really well in practice, but it's root being square error evaluation would look really bad. So an alternate method that actually captures this idea is to use a method called precision at top k. So let's say the user we, we look at the withheld ratings of the user, but only look at the the highest k ratings in the withheld set for a given user. And see whether our algorithm can correct a large fraction of those items. So the position at top k is just a percentage of the predictions that are the user's top k withheld ratings.