1
00:00:00,650 --> 00:00:02,930
Welcome back to Mining
of Massive Datasets.

2
00:00:02,930 --> 00:00:04,472
Today's topic is Recommender Systems.

3
00:00:04,472 --> 00:00:07,500
We're going to start with an overview
of recommendation systems, and

4
00:00:07,500 --> 00:00:08,166
why they are necessary.

5
00:00:08,166 --> 00:00:11,722
Today, we are going to look at the two
most common types of Recommender Systems.

6
00:00:11,722 --> 00:00:14,720
Content-Based Systems and
Collaborative Filtering.

7
00:00:14,720 --> 00:00:17,330
And finally we're going to look at
how to evaluate Recommender Systems,

8
00:00:17,330 --> 00:00:19,420
to make sure they're doing a good job.

9
00:00:19,420 --> 00:00:20,240
Let's start with an overview.

10
00:00:22,790 --> 00:00:26,490
Imagine any situation where
a user interacts with

11
00:00:26,490 --> 00:00:28,580
a really large catalog of items.

12
00:00:28,580 --> 00:00:30,910
Now these items could
be products at Amazon.

13
00:00:30,910 --> 00:00:32,530
They could be movies at Netflix.

14
00:00:32,530 --> 00:00:34,610
They could be music
from Pandora's catalog.

15
00:00:34,610 --> 00:00:37,480
Or they could be the news
items on Google News.

16
00:00:37,480 --> 00:00:40,180
What really matters,
is that there are tens of thousands, or

17
00:00:40,180 --> 00:00:41,950
hundreds of thousands,
or millions of items.

18
00:00:41,950 --> 00:00:43,330
A really large catalog.

19
00:00:43,330 --> 00:00:45,920
And the user is interacting
with this catalog.

20
00:00:45,920 --> 00:00:49,960
There's two ways in which a user can
interact with a large catalog of items.

21
00:00:49,960 --> 00:00:53,670
The first is such, user knows what
they're looking for, and they go and

22
00:00:53,670 --> 00:00:56,800
they search the catalog for
the precise item that they're looking for.

23
00:00:56,800 --> 00:00:59,280
Now when you have a very
large catalog of items,

24
00:00:59,280 --> 00:01:02,020
very often the user doesn't know
exactly what they're looking for.

25
00:01:02,020 --> 00:01:03,320
And this is where recommendations come in.

26
00:01:04,390 --> 00:01:08,490
The system recommends to the user certain
items that they think the user will be

27
00:01:08,490 --> 00:01:10,690
interested in,
based on what they know about the user.

28
00:01:11,820 --> 00:01:13,730
Now why do we really need
such recommendations?

29
00:01:15,650 --> 00:01:19,280
The key that made recommendations so
important and

30
00:01:19,280 --> 00:01:22,900
why a recommendation system developed so
much in the last ten or 20 years.

31
00:01:22,900 --> 00:01:26,280
If that he moved from an area
scarcity to an area of abundance.

32
00:01:26,280 --> 00:01:27,780
What do I mean by this?

33
00:01:27,780 --> 00:01:30,680
Imagine that you were out
shopping 20 years ago,

34
00:01:30,680 --> 00:01:32,460
and you'd go to a local retailer, and

35
00:01:32,460 --> 00:01:36,270
you'll find a certain number of products
on the shelves of the local retailer.

36
00:01:36,270 --> 00:01:40,180
Now, even in the really large retailer
like, like a Wal Mart, for instance.

37
00:01:40,180 --> 00:01:43,620
Shelf space is a key,
is a scarce commodity.

38
00:01:43,620 --> 00:01:47,300
It limits the number of items
that a retailer can carry.

39
00:01:47,300 --> 00:01:51,050
Shelf space is expensive,
because it involves real estate costs.

40
00:01:51,050 --> 00:01:54,610
And, therefore, a retailer can carry
only a certain number of products.

41
00:01:54,610 --> 00:01:58,390
Now, a similar situation applies in
the case of, for example, TV networks.

42
00:01:58,390 --> 00:02:01,190
A TV network can carry only so
many shows, because there's only so

43
00:02:01,190 --> 00:02:02,830
many hours in a day.

44
00:02:02,830 --> 00:02:03,440
And there are only so

45
00:02:03,440 --> 00:02:06,740
many movie theaters, so they can only ser,
screen a certain number of movies.

46
00:02:08,180 --> 00:02:11,290
Now once the internet was developed,
things changed.

47
00:02:11,290 --> 00:02:16,340
The web enables zero-cost dissemination
of information about products.

48
00:02:16,340 --> 00:02:20,380
And what this means, is that we can have
many more products than ever before.

49
00:02:20,380 --> 00:02:24,090
There is no shelf space limitation
on the number of products.

50
00:02:24,090 --> 00:02:26,980
That's why the number of
products on Amazon is much,

51
00:02:26,980 --> 00:02:30,390
much more than the number of products
available at any physical retailer.

52
00:02:32,520 --> 00:02:36,300
The number of you know, movies available
on Netflix is more than the number of

53
00:02:36,300 --> 00:02:39,010
movies that have been available,
available at Blockbuster and so on.

54
00:02:41,080 --> 00:02:44,580
This near-zero-cost dissemination
of information gives rise to

55
00:02:44,580 --> 00:02:47,400
a phenomenon that's called
the long tail phenomenon.

56
00:02:47,400 --> 00:02:48,570
Let's examine what this is.

57
00:02:50,010 --> 00:02:54,570
Now imagine a graph, where on X-axis
we've taken the items in the catalog.

58
00:02:54,570 --> 00:02:59,080
Remember, items might be books, or
music, or video, or news articles.

59
00:02:59,080 --> 00:03:01,530
And we've ranked these
items by popularity.

60
00:03:01,530 --> 00:03:03,920
So the most popular items are on the left,
and

61
00:03:03,920 --> 00:03:07,550
as they move towards the right,
the items become less and less popular.

62
00:03:07,550 --> 00:03:08,850
What do I mean by popular?

63
00:03:08,850 --> 00:03:12,240
Well, I mean the number of times
the item is purchased in a week.

64
00:03:12,240 --> 00:03:15,620
Or the number of times a movie is
viewed in a week, or a month, or

65
00:03:15,620 --> 00:03:17,890
some, some fixed time period.

66
00:03:17,890 --> 00:03:21,570
Now, on the Y axis, you have the actual
popularity, which in this case I've

67
00:03:21,570 --> 00:03:24,850
shown as the number of purchases per week,
it could be number of views per week, or

68
00:03:24,850 --> 00:03:29,100
it could be number of, you know, plays
per month, for some music, and so on.

69
00:03:29,100 --> 00:03:32,950
So in general you have items ranked
by popularity along the X axis,

70
00:03:32,950 --> 00:03:35,359
and the popularity
itself along the Y axis.

71
00:03:37,110 --> 00:03:40,870
Now when you take items you know,
in a large catalog.

72
00:03:40,870 --> 00:03:42,250
And you rank them,

73
00:03:42,250 --> 00:03:46,390
and you plot them on this curve you
get a curve that looks like this.

74
00:03:48,860 --> 00:03:52,687
You can see that the score, you know,
has a very steep fall initially.

75
00:03:52,687 --> 00:03:55,480
the, the, you know, you have a really,
really, a few really,

76
00:03:55,480 --> 00:03:57,090
really popular items.

77
00:03:57,090 --> 00:03:59,430
And then as you move
towards the right as the,

78
00:03:59,430 --> 00:04:04,270
you know, as the item rank becomes greater
the popularity falls off very steeply.

79
00:04:04,270 --> 00:04:07,790
But at a certain point, you can see
that this popularity stops, you know,

80
00:04:07,790 --> 00:04:09,740
falls off less and less deeply.

81
00:04:09,740 --> 00:04:12,480
And, you know,
it quite reaches the X axis.

82
00:04:13,830 --> 00:04:16,860
The interesting thing here,
is that there is a cut off point.

83
00:04:16,860 --> 00:04:20,590
The you know, items that are less
popular than this cut off point.

84
00:04:20,590 --> 00:04:23,270
You know, might be purchased
perhaps just once a week.

85
00:04:23,270 --> 00:04:24,760
Or maybe once a month.

86
00:04:24,760 --> 00:04:28,620
If you're a physical retailer like
a Wal-Mart, it's not economic to

87
00:04:28,620 --> 00:04:33,470
stock this item, because the rent cost of
stocking the item is more than you make,

88
00:04:33,470 --> 00:04:35,280
when you sell the item.

89
00:04:35,280 --> 00:04:36,850
And therefore a retailer,

90
00:04:36,850 --> 00:04:41,250
any right thinking retailer doesn't
stock items that are unpopular.

91
00:04:41,250 --> 00:04:44,400
The, you know, they only stock the,
the head of the distribution.

92
00:04:44,400 --> 00:04:50,700
So there's this cutoff point that I show
on this graph here and items that are more

93
00:04:50,700 --> 00:04:55,040
popular then this, the, the more popular
items are available at a retail store.

94
00:04:55,040 --> 00:04:59,010
But the less popular items, the items that
are to the right of the cut off point,

95
00:04:59,010 --> 00:05:00,810
are not available at any retail store.

96
00:05:00,810 --> 00:05:01,960
They're only available online.

97
00:05:04,000 --> 00:05:06,880
Now this phenomenon applies to books,
to music, to movie,

98
00:05:06,880 --> 00:05:09,750
to videos to news articles for
example, there are only so

99
00:05:09,750 --> 00:05:13,740
many news articles in newspaper, but
when you go online you can see the rest of

100
00:05:13,740 --> 00:05:16,579
the news articles are less popular,
news articles that are off to the right.

101
00:05:19,140 --> 00:05:24,100
The piece of the curve, that is to
the the piece of the curve here that is to

102
00:05:24,100 --> 00:05:28,520
the right of this dividing line,
is called the long tail.

103
00:05:28,520 --> 00:05:31,290
These are the items that
are available only online.

104
00:05:31,290 --> 00:05:34,910
The interesting thing is, the,
is this area under the curve here.

105
00:05:34,910 --> 00:05:37,689
And you can see the area under
the curve here is quite significant.

106
00:05:38,980 --> 00:05:43,350
In fact, in some cases the area under the
curve on the right is about as large, or

107
00:05:43,350 --> 00:05:45,550
could be even larger than
the area of the curve,

108
00:05:45,550 --> 00:05:47,050
under the curve on the, on the left.

109
00:05:49,230 --> 00:05:52,930
So you have all these items that could
never be found in a physical store, but

110
00:05:52,930 --> 00:05:54,690
that can be only found online.

111
00:05:54,690 --> 00:05:57,480
But there are so many of them.

112
00:05:57,480 --> 00:06:02,060
That it's very hard for
any user to find all these items.

113
00:06:02,060 --> 00:06:04,330
Right, so
when you have the seed of abundance and

114
00:06:04,330 --> 00:06:08,970
you have so many items and
many of them are really found online.

115
00:06:08,970 --> 00:06:11,990
How, you know, how do you
introduce a user to all these new

116
00:06:11,990 --> 00:06:14,340
items they may have not otherwise find?

117
00:06:14,340 --> 00:06:16,960
When you have more choice like this,
when you have these millions and

118
00:06:16,960 --> 00:06:20,120
millions of items that are only available
online, you need a better way for

119
00:06:20,120 --> 00:06:22,360
the user to find all these items.

120
00:06:22,360 --> 00:06:24,440
The user doesn't even know
where to start looking, and

121
00:06:24,440 --> 00:06:26,310
that's where recommendation
engines come in.

122
00:06:28,430 --> 00:06:30,670
So recommendation engines
work in the case of many,

123
00:06:30,670 --> 00:06:34,210
many kinds of items books,
music, movies, news articles.

124
00:06:34,210 --> 00:06:36,600
Interestingly, they even
work in the case of people.

125
00:06:36,600 --> 00:06:38,640
For example, when you go to Facebook,
or LinkedIn, or

126
00:06:38,640 --> 00:06:42,800
Twitter, there are so many people that you
don't know who to follow or who to friend.

127
00:06:42,800 --> 00:06:44,070
And so Facebook, or LinkedIn, or

128
00:06:44,070 --> 00:06:48,630
Twitter make recommendations to you,
on the people you could follow or friend.

129
00:06:50,220 --> 00:06:53,640
I like this point with interesting
anecdote that shows you

130
00:06:53,640 --> 00:06:55,910
the power of a recommendation engine.

131
00:06:55,910 --> 00:07:01,850
Several years ago a book was
published called Touching the Void.

132
00:07:01,850 --> 00:07:02,980
It's a book about mountaineering.

133
00:07:02,980 --> 00:07:04,678
It's very, very good book.

134
00:07:04,678 --> 00:07:07,190
The book came out,
it didn't make much of a ripple.

135
00:07:07,190 --> 00:07:08,840
You know, few people bought the book.

136
00:07:08,840 --> 00:07:11,090
It got some decent reviews, but
it never became a bestseller.

137
00:07:12,940 --> 00:07:15,140
And then a few years
after Touching the Void,

138
00:07:15,140 --> 00:07:18,710
a new book was published on
mountaineering called Into Thin Air.

139
00:07:18,710 --> 00:07:20,710
Now Into Thin Air picked up traction,

140
00:07:20,710 --> 00:07:23,170
and lots of people started
buying Into Thin Air.

141
00:07:25,180 --> 00:07:28,150
Amazon noticed that a few of
people who bought Into Thin Air,

142
00:07:28,150 --> 00:07:30,350
had also bought Touching the Void.

143
00:07:30,350 --> 00:07:34,660
So they started recommending Touching the
Void, to people who bought Into Thin Air.

144
00:07:34,660 --> 00:07:38,270
And low and behold, those people started
buying Touching the Void as well.

145
00:07:38,270 --> 00:07:41,820
The interesting point is,
this made Touching the Void a bestseller.

146
00:07:41,820 --> 00:07:45,280
In fact, it became a bigger
bestseller even than Into Thin Air,

147
00:07:45,280 --> 00:07:48,450
even though a few years ago,
it had sank without a trace.

148
00:07:48,450 --> 00:07:52,620
So this example should show you
the power of recommendation systems.

149
00:07:52,620 --> 00:07:56,650
There are these items, these sort of
gems like Touching the Void you know,

150
00:07:56,650 --> 00:07:59,980
that people don't know because
they don't know to look for them.

151
00:07:59,980 --> 00:08:04,230
But a good recommendation system can
expose people to these hidden gems,

152
00:08:04,230 --> 00:08:05,720
that they wouldn't have
known about otherwise.

153
00:08:06,970 --> 00:08:09,390
So let's look at types of
recommendation systems.

154
00:08:10,600 --> 00:08:14,660
The simplest and the oldest kind
of recommendation is editorial or

155
00:08:14,660 --> 00:08:16,220
hand curated.

156
00:08:16,220 --> 00:08:19,070
You might find a list of favorites for
example when you go into your

157
00:08:19,070 --> 00:08:22,360
favorite neighborhood book store,
you might find staff picks.

158
00:08:22,360 --> 00:08:25,030
Certain marked off as staff picks, right?

159
00:08:25,030 --> 00:08:29,360
And these are editorial triangulated,
and on certain websites you'll

160
00:08:29,360 --> 00:08:33,510
see a list of staff favorites or
a list of essential items.

161
00:08:33,510 --> 00:08:36,200
These are essentially built by hand.

162
00:08:36,200 --> 00:08:40,310
And another place where you'll see these
editorial recommendations is often on

163
00:08:40,310 --> 00:08:41,820
the homepages of websites.

164
00:08:41,820 --> 00:08:46,940
For example if you go to
the the homepage of most popular

165
00:08:46,940 --> 00:08:51,940
websites including product websites
you'll see editorial picks.

166
00:08:51,940 --> 00:08:55,060
These are products that have been picked
by the editorial staff to feature

167
00:08:55,060 --> 00:08:55,590
on the home page.

168
00:08:55,590 --> 00:09:00,910
The drawback with editorial on
hand curated recommendation,

169
00:09:00,910 --> 00:09:05,740
is that the, it's done entirely by you
know, by the staff of the website, and

170
00:09:05,740 --> 00:09:08,270
there's no input from
the users of the site.

171
00:09:08,270 --> 00:09:10,360
So, when you go beyond
editorial recommendations,

172
00:09:10,360 --> 00:09:13,640
the next simple thing that you
can do is simple aggregates.

173
00:09:13,640 --> 00:09:17,560
On many websites,
you'll see lists of top ten, or

174
00:09:17,560 --> 00:09:21,020
most popular, or most recent for example,

175
00:09:21,020 --> 00:09:25,690
if you go to YouTube you can see the most
popular videos, for instance, right.

176
00:09:25,690 --> 00:09:28,590
So these are simple
aggregates which sort of

177
00:09:28,590 --> 00:09:32,650
take into account user activity to
make recommendations to other users.

178
00:09:32,650 --> 00:09:37,810
But these recommendations don't depend
on the user they only depend on,

179
00:09:37,810 --> 00:09:40,430
you know, the, the aggregate
activity of a lot of other users.

180
00:09:42,060 --> 00:09:44,990
The third and most interesting
kind of recommendation to us

181
00:09:44,990 --> 00:09:48,960
is recommendations that
are tailored to individual users.

182
00:09:48,960 --> 00:09:53,300
Right for example, book recommendations
tailored to your taste, or

183
00:09:53,300 --> 00:09:56,700
movie recommendations based on
the movies that you watched previously.

184
00:09:56,700 --> 00:09:59,970
Or music recommendation based
on your music interests.

185
00:09:59,970 --> 00:10:05,010
And this is our focus here recommendations
that are tailored to individual users.

186
00:10:09,790 --> 00:10:11,448
So lets look at a formal model.

187
00:10:11,448 --> 00:10:14,860
Let C be a set of customers and
S a set of items.

188
00:10:16,620 --> 00:10:21,410
We will create a function called
the utility function or a utility matrix.

189
00:10:21,410 --> 00:10:26,560
The utility function is a function
that looks at every pair of

190
00:10:26,560 --> 00:10:30,320
customer and item, and
maps it to a rating.

191
00:10:31,760 --> 00:10:32,570
Okay.

192
00:10:32,570 --> 00:10:34,980
R in this case is a set of ratings.

193
00:10:34,980 --> 00:10:40,270
And for example R could be a star
rating from one star to five star.

194
00:10:40,270 --> 00:10:43,595
R could be a number between zero and ten.

195
00:10:44,840 --> 00:10:48,400
In general,
R is a totally ordered set so that,

196
00:10:48,400 --> 00:10:52,470
you know, a lower value indicates that
a user liked the product less, and

197
00:10:52,470 --> 00:10:55,030
a higher value indicates that
a user liked the product more.

198
00:10:56,940 --> 00:10:58,870
Let's look at an example
of a utility matrix.

199
00:11:00,760 --> 00:11:03,220
On the top we have four movies here.

200
00:11:03,220 --> 00:11:06,250
Avatar, Lord of the Rings, Matrix,
and Pirates of the Caribbean.

201
00:11:06,250 --> 00:11:08,740
And down here we have four users Alice,
Bob, Carol, and David.

202
00:11:09,800 --> 00:11:15,150
And the utility matrix gives you rating
for certain movies and certain users.

203
00:11:15,150 --> 00:11:18,370
For example Alice has rated Avatar and

204
00:11:18,370 --> 00:11:22,210
Matrix, but not Lord of the Rings or
Pirates of the Caribbean.

205
00:11:22,210 --> 00:11:26,510
Whereas Carol has rated you know,
has rated the same two movies.

206
00:11:26,510 --> 00:11:27,970
Bob has rated Lord of the Rings and

207
00:11:27,970 --> 00:11:32,550
Pirates but
hasn't rated Avatar or, or Matrix.

208
00:11:32,550 --> 00:11:35,840
Now it could be that these users
have not scene these movies or

209
00:11:35,840 --> 00:11:39,199
it could be that they've seen the movies,
but not bother to rate them.

210
00:11:40,560 --> 00:11:44,570
So in general, usually a matrix
like this is going to be sparse.

211
00:11:44,570 --> 00:11:47,160
You know, most of the users
haven't seen most of the movies.

212
00:11:48,200 --> 00:11:53,010
And there are going to be values in some
of the you know, some of the locations.

213
00:11:53,010 --> 00:11:57,500
The problem in recommendation systems
is to figure out these unknown values.

214
00:11:57,500 --> 00:12:01,870
For example you've seen that
Alice is rated Avatar and

215
00:12:01,870 --> 00:12:04,820
Matrix, but
hasn't rated Lord Of The Rings.

216
00:12:04,820 --> 00:12:10,390
So the question is can we figure
out what Alice's rating for

217
00:12:10,390 --> 00:12:14,680
Lord Of The Rings will be,
based on her other ratings?

218
00:12:14,680 --> 00:12:17,270
Can you figure out whether
she likes Pirates or not?

219
00:12:18,600 --> 00:12:20,290
Right.
So this is the key problem for

220
00:12:20,290 --> 00:12:21,490
recommended systems.

221
00:12:21,490 --> 00:12:26,510
Once we find out for each user certain
movies that they would have rated highly,

222
00:12:26,510 --> 00:12:29,490
or the system thinks they
might have rated highly,

223
00:12:29,490 --> 00:12:31,750
then we can recommend those
movies to those users.

224
00:12:31,750 --> 00:12:36,930
So there are three key problems in
the space of recommended systems.

225
00:12:36,930 --> 00:12:41,050
The first is gathering the known
ratings in the ratings in the matrix.

226
00:12:41,050 --> 00:12:44,391
Now in the previous slide when you
look at the utility matrix it was

227
00:12:44,391 --> 00:12:46,207
already filled in with certain values.

228
00:12:46,207 --> 00:12:48,834
But how do you get,
gather those values in the first place?

229
00:12:48,834 --> 00:12:52,010
So that's the key that's the first
problem that you need to tackle.

230
00:12:54,080 --> 00:12:57,819
The second problem is to extrapolate
unknown ratings from known ratings.

231
00:12:59,770 --> 00:13:03,430
But we're mainly interested
in the high unknown ratings.

232
00:13:03,430 --> 00:13:06,750
We are interested in those ratings
where a user would've given

233
00:13:06,750 --> 00:13:07,730
a high rating to a movie.

234
00:13:07,730 --> 00:13:11,430
We are not interested in the average,
or the low ratings because they're never

235
00:13:11,430 --> 00:13:13,020
going to recommend those
movies to the user.

236
00:13:15,660 --> 00:13:20,180
And finally, the third key problem
is evaluating extrapolation methods.

237
00:13:20,180 --> 00:13:23,240
Once you have a recommendation system
that can extrapolate unknown ratings from

238
00:13:23,240 --> 00:13:27,070
known ratings, how do you know
the recommended system is doing well?

239
00:13:27,070 --> 00:13:30,162
This is where evaluation
methodologies come in.

240
00:13:30,162 --> 00:13:34,730
Let's start with the first problem,
that of gathering ratings.

241
00:13:36,010 --> 00:13:36,580
The first and

242
00:13:36,580 --> 00:13:41,510
simplest way of gathering ratings is,
is what I call explicit patterns.

243
00:13:41,510 --> 00:13:43,200
Simply ask people to rate items.

244
00:13:45,520 --> 00:13:47,100
Now this method is good because the,

245
00:13:47,100 --> 00:13:49,768
you know you're asking people
to directly rate items.

246
00:13:49,768 --> 00:13:51,640
And you're going to get, you know, for ex,

247
00:13:51,640 --> 00:13:54,630
and you can decide on what scale
people are going to rate items.

248
00:13:54,630 --> 00:13:58,720
For example, you can say and ask for
ratings on a one to five star scale.

249
00:13:58,720 --> 00:14:01,750
Or you can ask people to rate
on a scale from zero to ten.

250
00:14:02,920 --> 00:14:05,510
or, or we can just ask people to
say whether they like an item or

251
00:14:05,510 --> 00:14:07,100
did not like it.

252
00:14:07,100 --> 00:14:11,400
So the exclusive method had
the advantage of simplicity and

253
00:14:11,400 --> 00:14:13,470
of getting direct responses from users.

254
00:14:14,700 --> 00:14:16,260
The problem, though,
is that it doesn't scale.

255
00:14:18,010 --> 00:14:22,180
Only a small fraction of users who viewed
a movie or listened to a, you know, piece

256
00:14:22,180 --> 00:14:25,730
of music or bought a product are actually
bother to leave a rating or review.

257
00:14:26,880 --> 00:14:30,050
Most users don't actually leave ratings or
reviews.

258
00:14:30,050 --> 00:14:33,520
So while the data that's explicitly
gathered's excellent data.

259
00:14:34,670 --> 00:14:38,850
It it's not sufficient in most cases for
recommendations,

260
00:14:38,850 --> 00:14:42,009
because only a small fraction of users
actually leave ratings and reviews.

261
00:14:43,280 --> 00:14:48,400
Since explicit rating don't scale,
a lot of sites use implicit ratings.

262
00:14:48,400 --> 00:14:52,660
Now the idea behind implicit ratings is
to learn ratings from other user action.

263
00:14:52,660 --> 00:14:57,180
For example an online
shopping website might

264
00:14:58,510 --> 00:15:01,029
have a rule that a purchase
implies a high rating.

265
00:15:03,180 --> 00:15:05,130
Now the nice thing about implicit ratings,

266
00:15:05,130 --> 00:15:09,080
is that they're much more
scalable than explicit ratings.

267
00:15:09,080 --> 00:15:11,990
Because the user doesn't have
to explicitly rate an item, and

268
00:15:11,990 --> 00:15:15,160
there are way more other
actions than there are ratings.

269
00:15:16,360 --> 00:15:20,310
The problem though is that it's
very hard using implicit ratings to

270
00:15:20,310 --> 00:15:22,230
learn low ratings.

271
00:15:22,230 --> 00:15:24,410
It's quite easy to learn
high ratings because you,

272
00:15:24,410 --> 00:15:27,020
you might have it all that
points imply the high rating.

273
00:15:27,020 --> 00:15:34,036
But you, you can never learn a rating that
a user disliked a product implicitly.

274
00:15:34,036 --> 00:15:36,440
Impactors.

275
00:15:36,440 --> 00:15:37,660
Most recommender systems and

276
00:15:37,660 --> 00:15:41,950
most web sites, use a combination
of explicit and implicit ratings.

277
00:15:41,950 --> 00:15:45,030
Where explicitly ratings are available,
they use them.

278
00:15:45,030 --> 00:15:47,790
But they supplement them with
implicit ratings when needed.

279
00:15:51,730 --> 00:15:55,860
Let's move on now to the central
problem of Extrapolating Utilities.

280
00:15:55,860 --> 00:15:59,099
Or extrapolating unknown utilities
from known utility values.

281
00:16:02,240 --> 00:16:07,580
The key of central problem, that we have
this amount to extrapolate utilities,

282
00:16:07,580 --> 00:16:11,270
is that the matrix U,
the utility matrix is very, very sparse.

283
00:16:11,270 --> 00:16:13,500
Most people have not rated most items.

284
00:16:13,500 --> 00:16:16,800
And this introduces a slew of problems
that we'll come across shortly.

285
00:16:18,090 --> 00:16:20,810
The second problem we have
is a cold start problem.

286
00:16:20,810 --> 00:16:21,740
When you have a new item or

287
00:16:21,740 --> 00:16:26,870
a new user, the new item doesn't have any
ratings, and new users have no history.

288
00:16:26,870 --> 00:16:30,929
So, this is known as
the cold start problem, and

289
00:16:30,929 --> 00:16:34,906
how to tackle this problem as for
in due course.

290
00:16:34,906 --> 00:16:38,319
There are three approaches to
building recommender systems.

291
00:16:39,420 --> 00:16:41,900
The first is Content-based
recommendations.

292
00:16:41,900 --> 00:16:44,350
The second is Collaborative filtering.

293
00:16:44,350 --> 00:16:47,110
And the third is Latent
factor based modeling.

294
00:16:47,110 --> 00:16:49,410
Let's start with Content-based approaches.

