In the programming assignment for this module you will be implementing a content based filter in Lens kit. In another video, we'll be walking through exactly what you need to do for that assignment. But this video is going to walk through a simple but complete Lens kit recommender algorithm implementation To show you the moving pieces that are involved in building one. LensKit can, we understand, seem daunting and complicated. You have all these pieces, item scorers, models, builders, recommenders. And where to start, can, can see difficult to figure out sometimes. This video is going to walk through a minimal, but complete, LensKit recommender algorithm implementation. To show you the basic pieces that you need to look at if you're building a recommender algorithm. This code is the code you get if you run the, if you create a project using the fancy archetype. We use the archetype to make sure that our IDEs were working in the initial setup videos. You won't need to use archtypes for this class. We'll give you zip files that contain the projects starting points. But, the archtypes are there for your future use. Also the, the code that I'll be walking through here is also available as a download from the project page for the programming assignment for this module. So, the main components of a LensKit algorithm implementation typically are 3 classes, and then the other ones that they require. The ItemScorer is the heart of most algorithms. An item, the ItemScorer interface in Lenskit is an interface that provides user personalized scores for items. Most of the time, if you're creating a new algorithm on Lenskit you're creating a new ItemScrorer. Our default top end recommender will generate recommendations using that scorer. Our default rating predictor was use those scores and clamp them to the rating range to output rating prediction. But the ItemScorers were the heart of the logic for a lot of common recommender algorithms lie. Computing these user personalize scores for items that can be used to rank recommendations or to generate rating prediction. So most of the time in LensKit if you want to say I'm going to use this algorithm, really what you tell LensKit is, use this item scorer. Item scorers often use some form of a model, of precomputed data to aid their computations. This model can be an item similarity matrix. It can be just a table of, pre-computed mean ratings for each item. That's what the model and the code we'll look at in this video is. These models are separated from both, from the score so that they can be built, and then they can be serialized to disk, reloaded et ceterea. So you can have one machine that processes your data and builds a model You write the model to disk. You then load the model up in your web server in order to serve recommendations to your users. Then we typically have a separate class that does the model build, so the model build code is separated from the model itself. It keeps the code nice and cleanly architected. So lets look at how this code works out in practice. So I'm going to start by using the Maven Archtype to create a new project. And you can do this from your IDE, from Eclipse or from IntelliJ But I'm going to do it from the command line. And then I will import the project into intellij. So to generate a new project with an archetype, using the command line, the command is nvm, which is the main maven command, archetype colon generate. And this will scan the Maven repositories and my local repository for available arc types and then will prompt me for that and it gives me a list of 839 which I don't really feel like scrolling through so lets me just search so I can type lenskit And this will filter the list to show me the lenskit archetypes. I want the fancy one, which is number one, and then I want version 2.0.1. Now it'll ask me for a few properties for my new project. The groupId, which is how Maven names, part of how Maven names projects. I'll just do org dot grouplens. At lenskit dot demos and I'll call the project lenskit-example. And the version's fine. I don't really care about the version for this demonstration. And go ahead and reuse the group ID for the pro, for the package name. these are the settings I gave it. I like them. I'll tell it to create the project. Now it's created a project in the LensKit example directory. This is a maven project which most modern Java IDEs can import and interpret just fine. So now I'll start intellij. And I will import, or I will open, IntelliJ you don't actually have to do a, an official import. You can just open the Maven project, and IntelliJ will handle that just fine. So I'll go browse to where I created the archetype. And I will open the palm.xml file as a project. Intellij will scan it. And now. I have my project. To show that everything works, I'll run the project. So, to run it, I'll go to Run, and tell it I want to Run. And I don't have any Run configurations yet, so I'll edit configurations, and add a new one, and I'm going to add a Maven configuration. And I'm just going to run the entire lenskit valuation tool chain on it, which we'll discuss in more detail in another video. Show run the LensKit publish maven goal. And while it's running, we'll look at just a few things. First the pom dot xml is the maven project file that defines the project source code. It's dependencies and things. So, there's a variety of properties in here. We set the lens kit version that is using as a property just to make it easy to update the version. It says that this program depends on lens kit, the core, the K and M recommenders. We also use Junit for testing. And then it has a few Options, such as using Java 1.6 compatible code, and then setting up the LensKit evaluator. So, the project source code is like all maven projects. It implements a simple user item personalized mean. So, it takes the average rating for each item, and then it computes how much on average the user likes items with respect to the, the item averages. So, does the user tend to like items more or less than the average user? And it combines these 2 means. So you get a score that's just basically a linear regression to predict arraying across user, using user and item as the predictors. So, the, the, this class that extends the abstract item scorer which is a helper class that implements several of the Methods of the item scorer interface in terms of one, to reduce the work that you need to do in order to implement your own item scorer. We have a few fields in the item scorer. We have a user event DAO. And the dao, dao stands for data access object, and it's how lens kit provides access to the data underlying your recommender. The user event dao lets you look up events or ratings in the system by the user, by user ID. So you can, given a user ID, say, the user that you want to score for, you can look up all their ratings. We also have a user damping term to implement a slightly damped mean to keep from assuming that the user's one five star rating means that they like everything a lot. And then, we have the model. And in this class, this algorithm, the model is just a precomputed table of the mean and rating for each item. We then have a constructor that takes these three parameters and just saves them away into the fields. This constructors annotated with the add inject annotation. And that tells lens kit, and the dependency injector that lens kit uses, that this constructors supposed to be used to automatically build the item scorer. If you're familiar with tools that such as Guice or PicoContainers Spring, then you've seen dependency injection at work where the tool kit automatically instantiates your object using constructors like this. LensKit uses an injector called Draft which is, in many ways, very similar to Guice. So it will take this constructor, it will look at its' parameters to get the dependencies. In this case we depend on the Dao, the model, and then this damping parameter. It will automatically substantiate them and supply them to the scorer. Then I'll skip over this helper method for just a minute, the score method is the heart of the item square and it takes a user ID and this is the user that we gen, that was personalized in the scores for and it takes A mutable sparse vector and this vector is bold and input and output parameter, so the vector will contain in its key domain which is the set of valid ID. There are IDs that that vector can contain scores for. All of the items that the caller wants score for. And the Java of the score method If it's null, that means thatt the system doesn't know what, or doesn't know anything about this useer. So, we just use an empty set of ratings for that user, which will give them a mean offset of zero. As we'll see in just a little bit. We then convert their. Ratings into a rating vector. And the rating vector is a, vector that maps item ID's to the user's ratings. And this, so this rating vector user history summarizer is responsible for doing that and it just ha, has a static method to make a raing vector from a profile. This deals with the case where, your data source, remembers multiple ratings. So if the user re-rates an item, an you keep track of all those ratings with time stamps, the make rating vector will automatically just use the most recent rating as the users rating. An it gives you one rating for each item the user has rated. We then call this helper method compute user offset, which computes the user's average offset from item mean. And, we'll talk about that in just a minute. With that offset though, it then does the score. It fills in the scores, and it uses bulk methods on the sparse factor to do this efficiently. First, it fills the entire vector with the global mean rating, and even if the user has no offset and we've never seen that item before, global mean is, is exactly what we want to to score that item with. We then add the offset, the item offset, for every item. The item offset is the Ave, the, the difference between the items average and the global average rating, we then add in the users mean offset to finish the scoring computation. Let's take a minute now and go to. A piece of paper to look exactly what this formula is. So the scoring function that we're doing here is the, this user-personalized mean that is, the, so the, the score for an item, or for a user u, item i, is equal to the global mean rating Plus the item score, plus the user score, or baseline, or the user score. I'm just calling these b's, because we often use these things as baseline recommendeders, so The, the mu is just the global mean. B sub I is the item offset, and b sub u is the user. Offset. And b sub i is equal to the sum over the users that have rated The item. Of that user's rating minus the global mean. All over the number of users, the users, number of users that have righted the item. The user baseline the user score is, basically the same thing for users. Over each item the user has rated, the user's rating, minus that item's baseline. And then on both of these we just put in a little damping term gamma to decrease the extremism of root, means based on very few rating So that formula is exactly what this sequence of vector operations implements. We fill everything with the global mean, we add in each items offset and this, the, this version of the add method takes a vector, takes all the items that are in both vectors and just adds the value of one to the other. And then we add in the mean offset. And with mutable sparse vectors, the, the vector itself is being modified. The, the vector that you're invoking a method on is being modified. And this pattern Which might be slightly familiar to those of you who have done Fortran programming using a tool like Blast, allows us to very efficitiently do many computations and just accumulate results in a vector. So the user offset computation, this helper method we skipped, if the, the ratings is empty, we just shortcut and say, no offset. If the user doesn't have any ratings, we'll just use item averages as, they're scores. Then what we do is create a mutable copy of the users rating factors. So this method can freely modify this copy without oh, having that trickle out into the collar. We then subtract the global means. We subtract the item offsets. And this means we now have a vector of user offsets from item means. And we then just compute the damped mean. We take the sum of the vector. We divide it by its size, plus this user damping term. And the user damping term is usually going to be small. Something like 5. So that, once the user has quite a few ratings. The damping terms are effectively irrelevant in the file computation. So, now that we've seen this let's look at the i, the model just a little bit. The model for this is very simple. It consist of a global mean and a vector of item offsets from than global means, so that global mean plus item offset is the item's average rating. Now, this class does not have an at inject constructor, because we use what we call a, what's called a provider to build it. So this default provider annotation tell the dependency injector. To use item, when it needs a, an item mean model don't look for an add inject constructor. Instead instantiate an item mean model builder. Which might have its own add inject constructor and it's own dependencies. And call it get method to get one of these. And it's not necessary. We could implement all the logic in the item mean model. It's constructor, but I prefer to keep my constructor's simple and basically have them just copy data around and do some light computations. If there's heavy lifting to be done, I prefer to have that in a builder class to keep the code clean and the classes simple. There's also this sharable annotation and typically. Models have this annotation and what it means is that this annotation is serializable, and its thread safe, and it doesn't depend on any doubts and that means that lenskit can build it, it can write it to disk, it can load it back up and a web application or difference environment And it lets you move your models around between environments. So the model, the, this class is very simple. It just has the, the global mean. The items offsets getter. So that the score can get both of them. The interesting logic is in the model builder. The item mean model builder. And it implements the provider interface of item mean model. And it has a few, a couple of parameters- damping and it has the DAO. An event DAO, so we can get all the rate, all of the events, all the ratings and compute the average rating for each item and the global mean rating. So, it has an adject constructor that depends on both the event dow and this mean damping parameter. The event dow is annotated with this annotation, at transient. And at transient goes with, at sharable, and at transient really only shows up on constructor parameters for. Model builders. And it tells LensKit that the model builder will use the dao to build the model. But one the model was built, it doesn't not contain a reference to the dao. So, the dao is no longer needed once the model was built. And this enables some of LensKit's advanced features for automatically working with. Recommended configurations and it, there is not this condition that the model cannot use the DAO once it's build is not in forced. It's a promise that the model builder makes to LensKit and if that promise is not upheld things can break and if you'll remember, from the, the model. It doesn't have any reference to daos of any kind. So, the, the logic for the model builder is typically in the get method. And get is the only method defined by the provider interface. And it's just called to get whatever that provider is supposed to provide. A model builder. Provides by building the model, typically. So, in this, we're going to compute some averages. We're going to compute the, average rating for each item. And we're going to compute, global average. So, we, we, initialize a total, and a count, for the global average. We create a map that we're going to use to accumulate the sum of each item's ratings. Compute it's average. We then, initialize that, have a default return value of zero. And this is a fast util map which enables very fast operations on maps that don't have boxing overhead when you're working with primitives like longs and doubles. It's also the feature that lets you specify what's going to be returned if the key you look for isn't there. So, this just means, if a key's not there at 0, which is exactly what we want, if we haven't seen an item before, then it has no rating, so it's total rating is 0. We do the same thing for a vector, for a map of counts. We then get a cursor. Now, a cursor is basically an iterator. It also influence iterable that just returns itself, you can use a 4h loop with it. That is also, that needs to be closed when you're done so it could be backed by a file. It could be backed by a data base connection and it gives us a way to stream things into lens classes, in a type setting fashion. So it asks the DAO stream all events of type streaming. And we have a try finally so we close the rating cursor when we're done with it. We then loop over each rating. Now our rating contains a preference, and the preference is a user item rating triple. And if the preference is there, that means the user rated. If the preference is null, that means the user unrated the event. Most data sets you don't have unratings, but the one skit data model handles unrates as a general case. So we get the item ID, the value from the preference, we add, the value to the total, global total. We increase the global count, and then we just increase, The sum and the count for this particular item, taking advantage of the fact that they'll return to 0 if we've never seen the item before. After we've counted up and summed up everything, we can compute the global mean. And in all of these to keep things well defined we just use 0 if there's no values, there are no ratings what so ever. The global mean's easy, we then create a vector that's going to hold the item means, so we create a vector from the item rating count's keySet. so that's the, all the items we've seen. We're going to create a vector that has, that can hold those keys. We then iterate over every entry in the vector. We use fast iteration. Fast iteration is a pattern that LensKit borrowed from the fast util library, and what it says is that. What, what it means is that this vector entry is only going to be used within this loop and its not going to be kept a reference by any other code. So for each time to the loop we could modify and reuse the same vector enter, object since and vector entry is a fly weight we don't want to create too many of them if we don't have to. So fast is just a shorthand to tell the vector, get let's the loop, that the entries are not going to be references outside of a particular loop iteration. So feel free to mutate and return the same vector entry rather than. Creating a new one each time through. It makes the code more, it puts less pressure on the, on the garbage collector. So then we get the, the item id for that key, we also say that we want this state either, we want both set and unset, so this vector. It's going to have all the keys as valid keys, but none of them are set. So if we just iterated the vector, it's an empty vector. We're not going to get anything. But we say, vectorentry dot state dot either. To say, give us all the entries irrespective of whether or not they are set. We get the key. We then, get the. Item rating counts. Throw in the damping term. we get the sum. We also throw in the damping. And we put the damping in the, the numerator, as well. Algebra shows this to be equivalent to the, the formula that I showed you on the paper. And if there's. If the item count is positive, then we set the damp to mean minus the global mean. So it's an offset into, the vector. We then return an item mean model, that contains the mean, and an immutable copy of the vector using the freeze method. So that's a walk through of the main components of a LensKit, a working functional LensKit recommender algorithm. I hope you find it useful for getting your bearings as you work with LensKit and read the documentation.