1
00:00:00,580 --> 00:00:03,430
In the programming assignment for this
module you will

2
00:00:03,430 --> 00:00:07,870
be implementing a content based filter in
Lens kit.

3
00:00:07,870 --> 00:00:09,980
In another video, we'll be walking through
exactly

4
00:00:09,980 --> 00:00:12,470
what you need to do for that assignment.

5
00:00:12,470 --> 00:00:16,730
But this video is going to walk through a
simple but complete Lens kit

6
00:00:16,730 --> 00:00:19,820
recommender algorithm implementation To
show you the

7
00:00:19,820 --> 00:00:23,400
moving pieces that are involved in
building one.

8
00:00:23,400 --> 00:00:25,360
LensKit can, we understand, seem

9
00:00:25,360 --> 00:00:26,500
daunting and complicated.

10
00:00:26,500 --> 00:00:31,970
You have all these pieces, item scorers,
models, builders, recommenders.

11
00:00:31,970 --> 00:00:38,750
And where to start, can, can see difficult
to figure out sometimes.

12
00:00:38,750 --> 00:00:40,310
This video is going to walk through a

13
00:00:40,310 --> 00:00:44,390
minimal, but complete, LensKit recommender
algorithm implementation.

14
00:00:44,390 --> 00:00:46,410
To show you the basic pieces that you need

15
00:00:46,410 --> 00:00:48,900
to look at if you're building a
recommender algorithm.

16
00:00:49,970 --> 00:00:50,490
This code

17
00:00:50,490 --> 00:00:53,210
is the code you get if you run the,

18
00:00:53,210 --> 00:00:56,270
if you create a project using the fancy
archetype.

19
00:00:56,270 --> 00:00:59,320
We use the archetype to make sure that

20
00:00:59,320 --> 00:01:03,840
our IDEs were working in the initial setup
videos.

21
00:01:03,840 --> 00:01:06,120
You won't need to use archtypes for this
class.

22
00:01:06,120 --> 00:01:10,470
We'll give you zip files that contain the
projects starting points.

23
00:01:10,470 --> 00:01:13,220
But, the archtypes are there for your
future use.

24
00:01:13,220 --> 00:01:15,540
Also the, the code that I'll be walking

25
00:01:15,540 --> 00:01:18,910
through here is also available as a
download from

26
00:01:18,910 --> 00:01:24,550
the project page for the programming
assignment for this module.

27
00:01:24,550 --> 00:01:30,630
So, the main components of a LensKit
algorithm implementation typically

28
00:01:30,630 --> 00:01:34,390
are 3 classes, and then the other ones
that they require.

29
00:01:34,390 --> 00:01:37,560
The ItemScorer is the heart of most
algorithms.

30
00:01:37,560 --> 00:01:40,850
An item, the ItemScorer interface in
Lenskit is

31
00:01:40,850 --> 00:01:45,210
an interface that provides user
personalized scores for items.

32
00:01:46,330 --> 00:01:48,580
Most of the time, if you're creating a

33
00:01:48,580 --> 00:01:52,470
new algorithm on Lenskit you're creating a
new ItemScrorer.

34
00:01:52,470 --> 00:01:54,530
Our default top end recommender

35
00:01:54,530 --> 00:01:56,990
will generate recommendations using that
scorer.

36
00:01:56,990 --> 00:02:00,190
Our default rating predictor was use those
scores and

37
00:02:00,190 --> 00:02:04,420
clamp them to the rating range to output
rating prediction.

38
00:02:04,420 --> 00:02:05,910
But the ItemScorers were the

39
00:02:05,910 --> 00:02:10,610
heart of the logic for a lot of common
recommender algorithms lie.

40
00:02:10,610 --> 00:02:14,050
Computing these user personalize scores
for items that can

41
00:02:14,050 --> 00:02:19,270
be used to rank recommendations or to
generate rating prediction.

42
00:02:20,630 --> 00:02:23,490
So most of the time in LensKit if you want
to say I'm going

43
00:02:23,490 --> 00:02:28,420
to use this algorithm, really what you
tell LensKit is, use this item scorer.

44
00:02:30,690 --> 00:02:33,080
Item scorers often use some form of a

45
00:02:33,080 --> 00:02:37,200
model, of precomputed data to aid their
computations.

46
00:02:39,460 --> 00:02:43,010
This model can be an item similarity
matrix.

47
00:02:43,010 --> 00:02:48,670
It can be just a table of, pre-computed
mean ratings for each item.

48
00:02:48,670 --> 00:02:52,310
That's what the model and the code we'll
look at in this video is.

49
00:02:52,310 --> 00:02:56,680
These models are separated from both, from
the score so that they can

50
00:02:56,680 --> 00:03:00,860
be built, and then they can be serialized
to disk, reloaded et ceterea.

51
00:03:00,860 --> 00:03:04,500
So you can have one machine that processes
your data and

52
00:03:04,500 --> 00:03:06,930
builds a model You write the model to
disk.

53
00:03:06,930 --> 00:03:09,130
You then load the model up in your web

54
00:03:09,130 --> 00:03:12,560
server in order to serve recommendations
to your users.

55
00:03:12,560 --> 00:03:16,120
Then we typically have a separate class
that does the model

56
00:03:16,120 --> 00:03:20,130
build, so the model build code is
separated from the model itself.

57
00:03:20,130 --> 00:03:23,060
It keeps the code nice and cleanly
architected.

58
00:03:25,200 --> 00:03:30,670
So lets look at how this code works out in
practice.

59
00:03:30,670 --> 00:03:38,460
So I'm going to start by using the Maven
Archtype to create a new project.

60
00:03:38,460 --> 00:03:41,680
And you can do this from your IDE, from
Eclipse or

61
00:03:41,680 --> 00:03:45,650
from IntelliJ But I'm going to do it from
the command line.

62
00:03:45,650 --> 00:03:48,800
And then I will import the project into
intellij.

63
00:03:49,860 --> 00:03:55,540
So to generate a new project with an
archetype, using the command line,

64
00:03:55,540 --> 00:04:01,770
the command is nvm, which is the main
maven command, archetype colon generate.

65
00:04:03,350 --> 00:04:06,190
And this will scan the Maven repositories
and

66
00:04:06,190 --> 00:04:11,320
my local repository for available arc
types and then

67
00:04:11,320 --> 00:04:14,860
will prompt me for that and it gives me a
list of 839 which I don't really feel

68
00:04:14,860 --> 00:04:19,060
like scrolling through so lets me just
search so I can type

69
00:04:19,060 --> 00:04:23,640
lenskit And this will filter the list to
show me the lenskit archetypes.

70
00:04:23,640 --> 00:04:28,910
I want the fancy one, which is number one,
and then I want version 2.0.1.

71
00:04:28,910 --> 00:04:33,270
Now it'll ask me for a few properties for
my new project.

72
00:04:33,270 --> 00:04:39,220
The groupId, which is how Maven names,
part of how Maven names projects.

73
00:04:39,220 --> 00:04:39,940
I'll just

74
00:04:39,940 --> 00:04:47,021
do org dot grouplens.
At lenskit dot demos and I'll

75
00:04:47,021 --> 00:04:54,180
call the project lenskit-example.
And the version's fine.

76
00:04:54,180 --> 00:04:56,930
I don't really care about the version for
this demonstration.

77
00:04:56,930 --> 00:05:00,770
And go ahead and reuse the group ID for
the pro, for the package name.

78
00:05:02,450 --> 00:05:04,110
these are the settings I gave it.
I like them.

79
00:05:04,110 --> 00:05:06,580
I'll tell it to create the project.
Now

80
00:05:09,760 --> 00:05:12,870
it's created a project in the LensKit
example directory.

81
00:05:12,870 --> 00:05:15,370
This is a maven project which most modern

82
00:05:15,370 --> 00:05:19,240
Java IDEs can import and interpret just
fine.

83
00:05:19,240 --> 00:05:20,650
So now I'll start intellij.

84
00:05:29,660 --> 00:05:32,610
And I will import, or I will open,
IntelliJ

85
00:05:32,610 --> 00:05:35,250
you don't actually have to do a, an
official import.

86
00:05:35,250 --> 00:05:40,010
You can just open the Maven project, and
IntelliJ will handle that just fine.

87
00:05:40,010 --> 00:05:43,970
So I'll go browse to where I created the
archetype.

88
00:05:47,020 --> 00:05:52,540
And I will open the palm.xml file as a
project.

89
00:05:54,050 --> 00:05:55,170
Intellij will scan it.

90
00:05:57,920 --> 00:05:58,420
And now.

91
00:06:00,460 --> 00:06:01,540
I have my project.

92
00:06:03,100 --> 00:06:06,040
To show that everything works, I'll run
the project.

93
00:06:06,040 --> 00:06:11,440
So, to run it, I'll go to Run, and tell it
I want to Run.

94
00:06:11,440 --> 00:06:15,980
And I don't have any Run configurations
yet, so I'll edit configurations,

95
00:06:15,980 --> 00:06:20,420
and add a new one, and I'm going to add a
Maven configuration.

96
00:06:22,310 --> 00:06:26,260
And I'm just going to run the entire
lenskit

97
00:06:27,430 --> 00:06:31,730
valuation tool chain on it, which we'll
discuss in more detail in another video.

98
00:06:33,340 --> 00:06:36,899
Show run the LensKit publish maven goal.

99
00:06:43,350 --> 00:06:46,470
And while it's running, we'll look at just
a few things.

100
00:06:46,470 --> 00:06:53,170
First the pom dot xml is the maven project
file that defines the project source code.

101
00:06:53,170 --> 00:06:55,080
It's dependencies and things.

102
00:06:55,080 --> 00:06:58,130
So, there's a variety of properties in
here.

103
00:06:58,130 --> 00:07:00,040
We set the lens kit version that is using
as

104
00:07:00,040 --> 00:07:03,170
a property just to make it easy to update
the version.

105
00:07:03,170 --> 00:07:09,590
It says that this program depends on lens
kit, the core, the K and M recommenders.

106
00:07:09,590 --> 00:07:15,500
We also use Junit for testing.
And then it has a few Options,

107
00:07:15,500 --> 00:07:21,030
such as using Java 1.6 compatible code,

108
00:07:21,030 --> 00:07:26,560
and then setting up the LensKit evaluator.
So,

109
00:07:30,410 --> 00:07:33,560
the project source code is like all maven
projects.

110
00:07:54,670 --> 00:07:58,280
It implements a simple user item
personalized mean.

111
00:07:58,280 --> 00:08:03,530
So, it takes the average rating for each
item, and then it computes how

112
00:08:03,530 --> 00:08:08,400
much on average the user likes items with
respect to the, the item averages.

113
00:08:08,400 --> 00:08:13,820
So, does the user tend to like items more
or less than the average user?

114
00:08:15,380 --> 00:08:20,170
And it combines these 2 means.
So you get a score that's just

115
00:08:20,170 --> 00:08:24,020
basically a linear regression to predict

116
00:08:24,020 --> 00:08:27,590
arraying across user, using user and item
as the predictors.

117
00:08:29,410 --> 00:08:34,470
So, the, the, this class that extends the
abstract item scorer

118
00:08:34,470 --> 00:08:39,010
which is a helper class that implements
several of the

119
00:08:40,140 --> 00:08:45,170
Methods of the item scorer interface in
terms of one, to reduce the work that you

120
00:08:45,170 --> 00:08:47,660
need to do in order to implement your own
item scorer.

121
00:08:49,190 --> 00:08:51,590
We have a few fields in the item scorer.

122
00:08:51,590 --> 00:08:54,050
We have a user event DAO.

123
00:08:54,050 --> 00:08:57,970
And the dao, dao stands for data access
object, and it's

124
00:08:57,970 --> 00:09:02,460
how lens kit provides access to the data
underlying your recommender.

125
00:09:02,460 --> 00:09:05,260
The user event dao lets you look up events
or

126
00:09:05,260 --> 00:09:09,360
ratings in the system by the user, by user
ID.

127
00:09:09,360 --> 00:09:10,490
So you can, given a user

128
00:09:10,490 --> 00:09:14,379
ID, say, the user that you want to score
for, you can look up all their ratings.

129
00:09:15,710 --> 00:09:22,300
We also have a user damping term to
implement a slightly damped mean to keep

130
00:09:22,300 --> 00:09:25,070
from assuming that the user's one five
star

131
00:09:25,070 --> 00:09:28,530
rating means that they like everything a
lot.

132
00:09:29,880 --> 00:09:31,810
And then, we have the model.

133
00:09:31,810 --> 00:09:35,630
And in this class, this algorithm, the
model is just a precomputed

134
00:09:35,630 --> 00:09:37,760
table of the mean and rating for each
item.

135
00:09:38,950 --> 00:09:41,500
We then have a constructor that takes
these three

136
00:09:41,500 --> 00:09:45,580
parameters and just saves them away into
the fields.

137
00:09:45,580 --> 00:09:49,460
This constructors annotated with the add
inject annotation.

138
00:09:49,460 --> 00:09:53,800
And that tells lens kit, and the
dependency injector that lens kit uses,

139
00:09:53,800 --> 00:09:57,760
that this constructors supposed to be used
to automatically build the item scorer.

140
00:09:59,230 --> 00:10:00,940
If you're familiar with tools

141
00:10:00,940 --> 00:10:02,700
that such as Guice or PicoContainers

142
00:10:02,700 --> 00:10:05,880
Spring, then you've seen dependency
injection at

143
00:10:05,880 --> 00:10:08,800
work where the tool kit automatically

144
00:10:08,800 --> 00:10:12,942
instantiates your object using
constructors like this.

145
00:10:12,942 --> 00:10:15,770
LensKit uses an injector called Draft
which

146
00:10:15,770 --> 00:10:17,770
is, in many ways, very similar to Guice.

147
00:10:18,820 --> 00:10:22,910
So it will take this constructor, it will

148
00:10:22,910 --> 00:10:25,230
look at its' parameters to get the
dependencies.

149
00:10:25,230 --> 00:10:26,030
In this case we

150
00:10:26,030 --> 00:10:31,440
depend on the Dao, the model, and then
this damping parameter.

151
00:10:31,440 --> 00:10:36,660
It will automatically substantiate them
and supply them to the scorer.

152
00:10:38,290 --> 00:10:43,520
Then I'll skip over this helper method for
just a minute, the score method is

153
00:10:46,050 --> 00:10:51,337
the heart of the item square and it takes
a user ID and this is the user

154
00:10:51,337 --> 00:10:56,790
that we gen, that was personalized in the
scores for and it takes

155
00:10:56,790 --> 00:11:03,200
A mutable sparse vector and this vector is
bold and input and output parameter,

156
00:11:03,200 --> 00:11:09,050
so the vector will contain in its key
domain which is the set of valid ID.

157
00:11:09,050 --> 00:11:11,370
There are IDs that that vector can contain

158
00:11:11,370 --> 00:11:12,200
scores for.

159
00:11:13,220 --> 00:11:16,570
All of the items that the caller wants
score for.

160
00:11:16,570 --> 00:11:18,250
And the Java of the score method

161
00:11:38,680 --> 00:11:42,280
If it's null, that means thatt the system
doesn't

162
00:11:42,280 --> 00:11:46,510
know what, or doesn't know anything about
this useer.

163
00:11:46,510 --> 00:11:49,990
So, we just use an empty set of ratings
for

164
00:11:49,990 --> 00:11:54,090
that user, which will give them a mean
offset of zero.

165
00:11:54,090 --> 00:11:56,260
As we'll see in just a little bit.

166
00:11:56,260 --> 00:12:01,530
We then convert their.
Ratings into a rating vector.

167
00:12:01,530 --> 00:12:03,520
And the rating vector is a,

168
00:12:05,820 --> 00:12:09,590
vector that maps item ID's to the user's
ratings.

169
00:12:09,590 --> 00:12:13,040
And this, so this rating vector user

170
00:12:13,040 --> 00:12:15,640
history summarizer is responsible for
doing that and

171
00:12:15,640 --> 00:12:19,940
it just ha, has a static method to make a
raing vector from a profile.

172
00:12:19,940 --> 00:12:24,710
This deals with the case where, your data
source, remembers multiple ratings.

173
00:12:24,710 --> 00:12:27,600
So if the user re-rates an item, an you
keep

174
00:12:27,600 --> 00:12:30,870
track of all those ratings with time
stamps, the make rating

175
00:12:30,870 --> 00:12:34,600
vector will automatically just use the
most recent rating as the users rating.

176
00:12:34,600 --> 00:12:37,710
An it gives you one rating for each item
the user has rated.

177
00:12:40,310 --> 00:12:43,670
We then call this helper method compute
user offset,

178
00:12:43,670 --> 00:12:47,960
which computes the user's average offset
from item mean.

179
00:12:50,840 --> 00:12:53,400
And, we'll talk about that in just a
minute.

180
00:12:53,400 --> 00:12:56,640
With that offset though, it then does the
score.

181
00:12:56,640 --> 00:12:58,930
It fills in the scores, and it uses bulk

182
00:12:58,930 --> 00:13:01,830
methods on the sparse factor to do this
efficiently.

183
00:13:01,830 --> 00:13:04,870
First, it fills the entire vector with the
global

184
00:13:04,870 --> 00:13:08,170
mean rating, and even if the user has no
offset

185
00:13:08,170 --> 00:13:10,710
and we've never seen that item before,
global mean

186
00:13:10,710 --> 00:13:15,890
is, is exactly what we want to to score
that

187
00:13:15,890 --> 00:13:17,250
item with.

188
00:13:17,250 --> 00:13:23,200
We then add the offset, the item offset,
for every item.

189
00:13:23,200 --> 00:13:25,370
The item offset is the

190
00:13:27,600 --> 00:13:32,810
Ave, the, the difference between the items
average and the global average rating,

191
00:13:34,310 --> 00:13:39,530
we then add in the users mean offset to
finish the scoring computation.

192
00:13:42,500 --> 00:13:45,140
Let's take a minute now and go to.

193
00:13:45,140 --> 00:13:49,350
A piece of paper to look exactly what this
formula is.

194
00:13:49,350 --> 00:13:55,760
So the scoring function that we're doing
here is the, this user-personalized mean

195
00:13:55,760 --> 00:14:02,310
that is, the, so the, the score for an
item, or for a user u,

196
00:14:02,310 --> 00:14:09,180
item i, is equal to the global mean rating
Plus

197
00:14:10,870 --> 00:14:20,490
the item score, plus the user score, or
baseline, or the user score.

198
00:14:20,490 --> 00:14:23,000
I'm just calling these b's, because we
often use these things

199
00:14:23,000 --> 00:14:29,760
as baseline recommendeders, so The, the mu
is just the global mean.

200
00:14:38,020 --> 00:14:42,230
B sub I is the item offset,

201
00:14:44,790 --> 00:14:48,720
and b sub u is the user.
Offset.

202
00:14:51,070 --> 00:14:53,190
And b sub i

203
00:14:56,810 --> 00:15:00,070
is equal to the sum

204
00:15:02,500 --> 00:15:03,000
over

205
00:15:05,080 --> 00:15:11,280
the users that have rated The item.
Of that user's rating

206
00:15:13,370 --> 00:15:17,940
minus the global mean.
All over

207
00:15:23,170 --> 00:15:27,530
the number of users, the users, number of
users that have righted the item.

208
00:15:31,430 --> 00:15:35,250
The user baseline the user score is,

209
00:15:38,360 --> 00:15:44,880
basically the same thing for users.
Over

210
00:15:44,880 --> 00:15:52,188
each item

211
00:15:52,188 --> 00:15:56,840
the user has rated, the user's

212
00:15:56,840 --> 00:16:01,810
rating, minus that item's baseline.

213
00:16:10,350 --> 00:16:14,460
And then on both of these we just put in a
little damping term gamma

214
00:16:16,470 --> 00:16:24,079
to decrease the extremism of root, means
based on very few rating

215
00:16:25,590 --> 00:16:30,140
So that formula is exactly what this
sequence of vector operations implements.

216
00:16:30,140 --> 00:16:32,930
We fill everything with the global mean,
we add in

217
00:16:32,930 --> 00:16:36,760
each items offset and this, the, this
version of the

218
00:16:36,760 --> 00:16:39,150
add method takes a vector, takes all the
items that

219
00:16:39,150 --> 00:16:41,590
are in both vectors and just adds the
value of

220
00:16:41,590 --> 00:16:46,500
one to the other.
And then we add in the mean offset.

221
00:16:46,500 --> 00:16:52,230
And with mutable sparse vectors, the, the
vector itself is being modified.

222
00:16:52,230 --> 00:16:55,820
The, the vector that you're invoking a
method on is being modified.

223
00:16:55,820 --> 00:17:00,720
And this pattern Which might be slightly
familiar to those of you who have done

224
00:17:00,720 --> 00:17:04,990
Fortran programming using a tool like
Blast, allows

225
00:17:04,990 --> 00:17:07,380
us to very efficitiently do many
computations and

226
00:17:07,380 --> 00:17:09,220
just accumulate results in a vector.

227
00:17:10,590 --> 00:17:14,750
So the user offset computation, this
helper method we skipped, if

228
00:17:14,750 --> 00:17:18,670
the, the ratings is empty, we just
shortcut and say, no offset.

229
00:17:18,670 --> 00:17:20,310
If the user doesn't have any ratings,

230
00:17:20,310 --> 00:17:24,190
we'll just use item averages as, they're
scores.

231
00:17:25,350 --> 00:17:30,260
Then what we do is create a mutable copy
of the users rating factors.

232
00:17:30,260 --> 00:17:33,280
So this method can freely modify this copy

233
00:17:33,280 --> 00:17:38,010
without oh, having that trickle out into
the collar.

234
00:17:38,010 --> 00:17:40,780
We then subtract the global means.

235
00:17:40,780 --> 00:17:43,140
We subtract the item offsets.

236
00:17:43,140 --> 00:17:47,870
And this means we now have a vector of
user offsets from item means.

237
00:17:47,870 --> 00:17:49,720
And we then just compute the damped mean.

238
00:17:49,720 --> 00:17:50,940
We take the sum of the vector.

239
00:17:50,940 --> 00:17:53,570
We divide it by its size, plus this user
damping term.

240
00:17:53,570 --> 00:17:57,230
And the user damping term is usually
going to be small.

241
00:17:57,230 --> 00:17:58,270
Something like 5.

242
00:17:58,270 --> 00:18:00,100
So that, once the user has quite a few
ratings.

243
00:18:01,190 --> 00:18:04,240
The damping terms are effectively
irrelevant in the file computation.

244
00:18:05,780 --> 00:18:11,490
So, now that we've seen this let's look at
the i, the model just a little bit.

245
00:18:11,490 --> 00:18:13,940
The model for this is very simple.

246
00:18:13,940 --> 00:18:19,790
It consist of a global mean and a vector
of item offsets from than

247
00:18:19,790 --> 00:18:23,439
global means, so that global mean plus
item offset is the item's average rating.

248
00:18:25,560 --> 00:18:29,340
Now, this class does not have an at inject
constructor,

249
00:18:30,930 --> 00:18:36,030
because we use what we call a, what's
called a provider to build it.

250
00:18:36,030 --> 00:18:39,810
So this default provider annotation tell
the dependency injector.

251
00:18:41,430 --> 00:18:44,420
To use item, when it needs a, an item

252
00:18:44,420 --> 00:18:46,680
mean model don't look for an add inject
constructor.

253
00:18:46,680 --> 00:18:49,470
Instead instantiate an item mean model
builder.

254
00:18:49,470 --> 00:18:50,870
Which might have its own add inject

255
00:18:50,870 --> 00:18:53,160
constructor and it's own dependencies.

256
00:18:53,160 --> 00:18:55,390
And call it get method to get one of
these.

257
00:18:55,390 --> 00:18:56,440
And it's not necessary.

258
00:18:56,440 --> 00:19:00,199
We could implement all the logic in the
item mean model.

259
00:19:01,530 --> 00:19:06,070
It's constructor, but I prefer to keep my
constructor's simple and

260
00:19:06,070 --> 00:19:10,550
basically have them just copy data around
and do some light computations.

261
00:19:10,550 --> 00:19:13,100
If there's heavy lifting to be done, I
prefer to have that

262
00:19:13,100 --> 00:19:17,230
in a builder class to keep the code clean
and the classes simple.

263
00:19:17,230 --> 00:19:22,000
There's also this sharable annotation and
typically.

264
00:19:22,000 --> 00:19:27,660
Models have this annotation and what it
means is that this annotation

265
00:19:27,660 --> 00:19:32,100
is serializable, and its thread safe, and
it doesn't depend on any

266
00:19:32,100 --> 00:19:37,260
doubts and that means that lenskit can
build

267
00:19:37,260 --> 00:19:42,440
it, it can write it to disk, it can load
it back up and

268
00:19:42,440 --> 00:19:45,180
a web application or difference
environment And it

269
00:19:45,180 --> 00:19:49,190
lets you move your models around between
environments.

270
00:19:50,980 --> 00:19:53,990
So the model, the, this class is very
simple.

271
00:19:53,990 --> 00:19:57,250
It just has the, the global mean.
The items offsets getter.

272
00:19:57,250 --> 00:19:59,780
So that the score can get both of them.

273
00:20:00,900 --> 00:20:03,380
The interesting logic is in the model
builder.

274
00:20:03,380 --> 00:20:07,850
The item mean model builder.
And it implements the provider interface

275
00:20:07,850 --> 00:20:09,000
of item mean model.

276
00:20:10,130 --> 00:20:17,420
And it has a few, a couple of parameters-
damping and it has the DAO.

277
00:20:17,420 --> 00:20:22,090
An event DAO, so we can get all the rate,
all of the events, all the ratings

278
00:20:22,090 --> 00:20:27,639
and compute the average rating for each
item and the global mean rating.

279
00:20:28,880 --> 00:20:33,900
So, it has an adject constructor that
depends on both the event dow and

280
00:20:33,900 --> 00:20:37,910
this mean damping parameter.
The event dow is annotated with this

281
00:20:37,910 --> 00:20:42,580
annotation, at transient.
And at transient goes with, at sharable,

282
00:20:42,580 --> 00:20:49,770
and at transient really only shows up on
constructor parameters for.

283
00:20:49,770 --> 00:20:51,500
Model builders.

284
00:20:51,500 --> 00:20:54,490
And it tells LensKit that the model
builder

285
00:20:54,490 --> 00:20:56,670
will use the dao to build the model.

286
00:20:57,710 --> 00:20:58,990
But one the

287
00:20:58,990 --> 00:21:02,590
model was built, it doesn't not contain a
reference to the dao.

288
00:21:02,590 --> 00:21:05,000
So, the dao is no longer needed once the
model was built.

289
00:21:06,790 --> 00:21:08,340
And this enables some of LensKit's

290
00:21:08,340 --> 00:21:11,758
advanced features for automatically
working with.

291
00:21:11,758 --> 00:21:17,930
Recommended configurations and it, there
is not

292
00:21:17,930 --> 00:21:24,385
this condition that the model cannot use
the DAO once it's build is not in forced.

293
00:21:24,385 --> 00:21:28,117
It's a promise that the model builder
makes to LensKit

294
00:21:28,117 --> 00:21:31,790
and if that promise is not upheld things
can break

295
00:21:34,120 --> 00:21:36,920
and if you'll remember, from the, the
model.

296
00:21:36,920 --> 00:21:39,280
It doesn't have any reference to daos of
any kind.

297
00:21:41,240 --> 00:21:47,230
So, the, the logic for the model builder
is typically in the get method.

298
00:21:47,230 --> 00:21:50,390
And get is the only method defined by the
provider interface.

299
00:21:50,390 --> 00:21:53,860
And it's just called to get whatever that
provider is supposed to provide.

300
00:21:53,860 --> 00:21:55,450
A model builder.

301
00:21:55,450 --> 00:21:57,990
Provides by building the model, typically.

302
00:21:59,360 --> 00:22:02,520
So, in this, we're going to compute some
averages.

303
00:22:02,520 --> 00:22:07,880
We're going to compute the, average rating
for each item.

304
00:22:07,880 --> 00:22:10,800
And we're going to compute, global
average.

305
00:22:10,800 --> 00:22:15,440
So, we, we, initialize a total, and a
count, for the global average.

306
00:22:15,440 --> 00:22:20,980
We create a map that we're going to use to
accumulate the sum of each item's ratings.

307
00:22:20,980 --> 00:22:22,050
Compute it's average.

308
00:22:23,540 --> 00:22:24,910
We then,

309
00:22:24,910 --> 00:22:28,090
initialize that, have a default return
value of zero.

310
00:22:28,090 --> 00:22:33,360
And this is a fast util map which enables
very fast operations on maps

311
00:22:33,360 --> 00:22:35,570
that don't have boxing overhead when
you're

312
00:22:35,570 --> 00:22:37,350
working with primitives like longs and
doubles.

313
00:22:37,350 --> 00:22:40,250
It's also the feature that lets you
specify what's going

314
00:22:40,250 --> 00:22:44,180
to be returned if the key you look for
isn't there.

315
00:22:44,180 --> 00:22:46,920
So, this just means, if a key's not there
at 0, which

316
00:22:46,920 --> 00:22:50,630
is exactly what we want, if we haven't
seen an item before,

317
00:22:50,630 --> 00:22:53,470
then it has no rating, so it's total
rating is 0.

318
00:22:53,470 --> 00:22:57,590
We do the same thing for a vector, for a
map of counts.

319
00:22:57,590 --> 00:23:02,690
We then get a cursor.
Now, a cursor is basically an iterator.

320
00:23:02,690 --> 00:23:05,330
It also influence iterable that just
returns itself,

321
00:23:05,330 --> 00:23:07,220
you can use a 4h loop with it.

322
00:23:07,220 --> 00:23:09,710
That is also, that needs to be closed when

323
00:23:09,710 --> 00:23:12,080
you're done so it could be backed by a
file.

324
00:23:12,080 --> 00:23:16,010
It could be backed by a data base
connection and it gives us a way to stream

325
00:23:16,010 --> 00:23:20,700
things into lens classes, in a type
setting fashion.

326
00:23:20,700 --> 00:23:23,550
So it asks the DAO stream all events of
type streaming.

327
00:23:24,650 --> 00:23:29,130
And we have a try finally so we close the
rating cursor when we're done with it.

328
00:23:29,130 --> 00:23:31,490
We then loop over each rating.

329
00:23:31,490 --> 00:23:33,090
Now our rating contains a preference, and

330
00:23:33,090 --> 00:23:38,720
the preference is a user item rating
triple.

331
00:23:38,720 --> 00:23:42,110
And if the preference is there, that means
the user rated.

332
00:23:42,110 --> 00:23:46,380
If the preference is null, that means the
user unrated the event.

333
00:23:46,380 --> 00:23:49,620
Most data sets you don't have unratings,
but the

334
00:23:49,620 --> 00:23:53,090
one skit data model handles unrates as a
general case.

335
00:23:53,090 --> 00:23:54,670
So we get the item ID, the value from the

336
00:23:54,670 --> 00:23:59,350
preference, we add, the value to the
total, global total.

337
00:23:59,350 --> 00:24:01,680
We increase the global count, and then we
just

338
00:24:01,680 --> 00:24:07,480
increase, The sum and the count for this
particular item,

339
00:24:07,480 --> 00:24:09,430
taking advantage of the fact that they'll
return

340
00:24:09,430 --> 00:24:12,110
to 0 if we've never seen the item before.

341
00:24:14,790 --> 00:24:17,710
After we've counted up and summed up

342
00:24:17,710 --> 00:24:20,200
everything, we can compute the global
mean.

343
00:24:21,510 --> 00:24:24,480
And in all of these to keep things well
defined we just

344
00:24:24,480 --> 00:24:28,870
use 0 if there's no values, there are no
ratings what so ever.

345
00:24:30,460 --> 00:24:34,460
The global mean's easy, we then create a
vector that's going to hold

346
00:24:34,460 --> 00:24:39,130
the item means, so we create a vector from
the item rating count's keySet.

347
00:24:40,430 --> 00:24:42,840
so that's the, all the items we've seen.

348
00:24:42,840 --> 00:24:46,080
We're going to create a vector that has,
that can hold those keys.

349
00:24:47,320 --> 00:24:50,900
We then iterate over every entry in the
vector.

350
00:24:50,900 --> 00:24:52,200
We use fast iteration.

351
00:24:52,200 --> 00:24:55,520
Fast iteration is a pattern that LensKit
borrowed from

352
00:24:55,520 --> 00:25:00,160
the fast util library, and what it says is
that.

353
00:25:01,300 --> 00:25:05,630
What, what it means is that this vector
entry is only going to be used

354
00:25:05,630 --> 00:25:10,880
within this loop and its not going to be
kept a reference by any other code.

355
00:25:10,880 --> 00:25:14,410
So for each time to the loop we could

356
00:25:14,410 --> 00:25:18,614
modify and reuse the same vector enter,
object since and

357
00:25:18,614 --> 00:25:21,140
vector entry is a fly weight we don't want
to

358
00:25:21,140 --> 00:25:22,950
create too many of them if we don't have
to.

359
00:25:22,950 --> 00:25:26,030
So fast is just a shorthand to tell

360
00:25:26,030 --> 00:25:31,040
the vector, get let's the loop, that the
entries

361
00:25:31,040 --> 00:25:34,000
are not going to be references outside of
a particular loop iteration.

362
00:25:34,000 --> 00:25:40,300
So feel free to mutate and return the same
vector entry rather than.

363
00:25:40,300 --> 00:25:42,760
Creating a new one each time through.

364
00:25:42,760 --> 00:25:46,880
It makes the code more, it puts less
pressure on the, on the garbage collector.

365
00:25:48,470 --> 00:25:51,490
So then we get the, the item id for that
key, we also say

366
00:25:51,490 --> 00:25:56,610
that we want this state either, we want
both set and unset, so this vector.

367
00:25:56,610 --> 00:25:59,690
It's going to have all the keys as valid
keys, but none of them are set.

368
00:25:59,690 --> 00:26:01,760
So if we just iterated the vector, it's an
empty vector.

369
00:26:01,760 --> 00:26:02,790
We're not going to get anything.

370
00:26:03,970 --> 00:26:06,910
But we say, vectorentry dot state dot
either.

371
00:26:06,910 --> 00:26:09,370
To say, give us all the entries

372
00:26:09,370 --> 00:26:11,310
irrespective of whether or not they are
set.

373
00:26:12,610 --> 00:26:16,540
We get the key.
We then, get the.

374
00:26:16,540 --> 00:26:17,940
Item rating counts.

375
00:26:19,110 --> 00:26:21,870
Throw in the damping term.
we get the sum.

376
00:26:23,410 --> 00:26:26,220
We also throw in the damping.

377
00:26:26,220 --> 00:26:30,470
And we put the damping in the, the
numerator, as well.

378
00:26:30,470 --> 00:26:32,650
Algebra shows this to be equivalent to
the,

379
00:26:32,650 --> 00:26:34,740
the formula that I showed you on the
paper.

380
00:26:34,740 --> 00:26:35,520
And if there's.

381
00:26:36,990 --> 00:26:42,570
If the item count is positive, then we set
the damp to mean minus the global mean.

382
00:26:42,570 --> 00:26:49,030
So it's an offset into, the vector.
We then return an item mean model,

383
00:26:49,030 --> 00:26:51,390
that contains the mean, and an immutable

384
00:26:51,390 --> 00:26:54,300
copy of the vector using the freeze
method.

385
00:26:54,300 --> 00:26:57,130
So that's a walk through of the main
components

386
00:26:57,130 --> 00:27:02,140
of a LensKit, a working functional LensKit
recommender algorithm.

387
00:27:02,140 --> 00:27:04,260
I hope you find it useful for getting your

388
00:27:04,260 --> 00:27:07,750
bearings as you work with LensKit and read
the documentation.

