1
00:00:00,490 --> 00:00:02,620
Welcome back to Mining
of Massive Datasets.

2
00:00:02,620 --> 00:00:05,898
We're going to continue our lesson
on recommender systems by looking at

3
00:00:05,898 --> 00:00:10,770
content-based recommendation systems.

4
00:00:10,770 --> 00:00:15,030
The main idea behind content-based
recommendation systems is to recommend

5
00:00:15,030 --> 00:00:20,470
items to a customer x similar to previous
items rated highly by the same customer.

6
00:00:21,970 --> 00:00:26,230
For example, in, in example of
movies you might recommend movies

7
00:00:26,230 --> 00:00:31,490
with the same actor of actors,
director, genre and so on.

8
00:00:31,490 --> 00:00:33,160
In the case of websites, blogs or

9
00:00:33,160 --> 00:00:38,450
news, we might recommend articles
with similar content, similar topics.

10
00:00:38,450 --> 00:00:40,210
In the case of people recommendations,

11
00:00:40,210 --> 00:00:43,060
we might recommendation people with
many common friends to each other.

12
00:00:45,020 --> 00:00:45,910
So here's our plan of action.

13
00:00:45,910 --> 00:00:48,420
We're going to start with a user and

14
00:00:48,420 --> 00:00:53,620
find out a set of items the user likes
using both explicit and implicit data.

15
00:00:55,520 --> 00:00:59,740
For example, we might look at the items
that the user has rated highly and

16
00:00:59,740 --> 00:01:02,180
the set of items the user has purchased.

17
00:01:02,180 --> 00:01:05,180
And for eat of, each of those items,
we're going to build an item profile.

18
00:01:06,350 --> 00:01:09,020
An item profile is
a description of the item.

19
00:01:10,030 --> 00:01:13,400
For example, in this case,
we are dealing with geometric shapes.

20
00:01:13,400 --> 00:01:20,220
And let's say the user likes, they,
a red circle and, a red triangle.

21
00:01:20,220 --> 00:01:25,860
We might build item profiles that
say that the user likes red items.

22
00:01:25,860 --> 00:01:26,370
Right?

23
00:01:26,370 --> 00:01:29,150
Or they, or the user likes circles,
for instance.

24
00:01:32,320 --> 00:01:35,550
And from these item from
these item profiles,

25
00:01:35,550 --> 00:01:38,340
you're going to infer a user profile.

26
00:01:38,340 --> 00:01:41,620
The user profile infers the likes of

27
00:01:41,620 --> 00:01:44,810
the user from the profile
of items the user likes.

28
00:01:44,810 --> 00:01:47,630
Because the user here
likes a red circle and

29
00:01:47,630 --> 00:01:51,670
a red triangle we would infer that
the user likes the color red.

30
00:01:51,670 --> 00:01:52,900
They like circles.

31
00:01:52,900 --> 00:01:53,770
And they like triangles.

32
00:01:55,090 --> 00:01:59,740
Now once we have a profile of the user we
can then match that against a catalog and

33
00:01:59,740 --> 00:02:01,600
recommend our items to the user.

34
00:02:01,600 --> 00:02:04,870
So, let's say the catalog
has a bunch of items in it.

35
00:02:04,870 --> 00:02:07,640
Some of those items are red, so
we can recommend those to the user.

36
00:02:10,180 --> 00:02:13,830
So, let's, let's look at how
to build these item profiles.

37
00:02:13,830 --> 00:02:14,880
For each item,

38
00:02:14,880 --> 00:02:18,690
you want to create an item profile, which
we can then use to build user profile.

39
00:02:20,500 --> 00:02:23,680
So, the profile is a set
of features about the item.

40
00:02:23,680 --> 00:02:25,400
In the case of movies for instance.

41
00:02:25,400 --> 00:02:32,220
The item profile might include author
title actor director and so on.

42
00:02:32,220 --> 00:02:34,390
In the case of images and videos.

43
00:02:34,390 --> 00:02:36,710
We might use metadata and tags.

44
00:02:36,710 --> 00:02:41,240
In the case of people the item profile
might be a set of friends of the user.

45
00:02:41,240 --> 00:02:44,280
Given that the item profile
is a set of features,

46
00:02:44,280 --> 00:02:47,220
it's often convenient to
think of it as a vector.

47
00:02:47,220 --> 00:02:51,170
The vector could be either Boolean or real
valued, and there's one entry per feature.

48
00:02:52,400 --> 00:02:56,460
For example, in the case of movies,
the vector might be,

49
00:02:56,460 --> 00:02:59,010
the item profile might
be a Boolean vector.

50
00:02:59,010 --> 00:03:01,340
And there's a 0 or a 1.

51
00:03:01,340 --> 00:03:05,420
For each actor, director and so on,
depending on whether that actor or

52
00:03:05,420 --> 00:03:08,489
that director actually
participated in that movie.

53
00:03:13,427 --> 00:03:16,170
We'll look at the special case of text.

54
00:03:16,170 --> 00:03:18,480
For example,
we might be recommending news articles.

55
00:03:19,910 --> 00:03:22,379
Now, what's the item profile in this case.

56
00:03:23,630 --> 00:03:27,410
The simplest item profile in this case
is to pick the set of important words in

57
00:03:27,410 --> 00:03:28,540
the document or the item.

58
00:03:29,830 --> 00:03:31,540
How do you pick the important
words in the item?

59
00:03:31,540 --> 00:03:37,100
The usual heuristic that we get from text
mining is the technique called TF-IDF,

60
00:03:37,100 --> 00:03:40,120
or term frequency inverse
document frequency.

61
00:03:40,120 --> 00:03:44,450
Many of you may have come across TF-IDF in
the context of information retrieval but

62
00:03:44,450 --> 00:03:46,590
for those of you who have not
here's a quick refresher.

63
00:03:50,410 --> 00:03:52,890
Let's say we are looking at a document or

64
00:03:52,890 --> 00:03:58,290
item j, and is computing the score for
term or feature i.

65
00:03:59,390 --> 00:04:05,940
The term frequency TFIJ for
feature I in document J

66
00:04:05,940 --> 00:04:11,660
is just the number of times the feature J,
the feature I appears in the document J

67
00:04:11,660 --> 00:04:17,460
divided by the maximum number of time that
same feature appears in any document.

68
00:04:17,460 --> 00:04:21,820
For example, let's say the feature
is a certain word the word apple.

69
00:04:23,160 --> 00:04:28,530
And, int he document that we're looking
at, the word apple appears five times.

70
00:04:28,530 --> 00:04:32,260
But there's another document where
the word apple appears 23 times.

71
00:04:32,260 --> 00:04:35,710
And this is the maximum number of
times the word apple appears in

72
00:04:35,710 --> 00:04:37,060
any document at all.

73
00:04:37,060 --> 00:04:42,910
Then the term frequency, TF ij is,
it's five divided by 23.

74
00:04:42,910 --> 00:04:47,920
Now, I'm glossing over the fact that
we need to normalize TF to account for

75
00:04:47,920 --> 00:04:49,510
the fact that document
lengths are different.

76
00:04:49,510 --> 00:04:51,720
Let's just ignore that for the moment.

77
00:04:51,720 --> 00:04:52,670
Now, the,

78
00:04:52,670 --> 00:04:58,230
term frequency captures the number of
times, A term of years in a document.

79
00:04:58,230 --> 00:05:02,570
Intuitively, the more often,
term of years in a document,

80
00:05:02,570 --> 00:05:04,260
the more important a feature it is.

81
00:05:04,260 --> 00:05:10,920
For example, if a document mentions about
Apple five times, we weight Apple as more

82
00:05:10,920 --> 00:05:14,270
important in that document than in other
documents that just mentions it once.

83
00:05:16,430 --> 00:05:20,000
But how do you compare
the weight of different terms?

84
00:05:20,000 --> 00:05:25,520
For example, you know a, a rare word
appearing just a couple of times might

85
00:05:25,520 --> 00:05:30,070
more important than a more common word
like the appearing thousands of times.

86
00:05:30,070 --> 00:05:32,950
This is where the document
frequency comes in.

87
00:05:32,950 --> 00:05:37,020
Let n i be the number of
documents that mention the term i.

88
00:05:37,020 --> 00:05:40,670
And let n be the total number of
documents, in the whole system.

89
00:05:41,740 --> 00:05:48,210
The inverse document frequency for
the term i is obtained by dividing,

90
00:05:48,210 --> 00:05:52,610
N by n i, the number of documents
that mention the term i and

91
00:05:52,610 --> 00:05:55,590
then taking the logarithm of that,
of that, fraction.

92
00:05:57,580 --> 00:06:00,342
Notice that the more common a term,
the larger an i.

93
00:06:00,342 --> 00:06:05,960
And the larger an i, the lower the IDF.

94
00:06:07,750 --> 00:06:09,780
The IDF function ensures, you know,

95
00:06:09,780 --> 00:06:14,920
gives a lower weight to more common words
and a higher weight to rarer words.

96
00:06:16,190 --> 00:06:18,250
So if you put these two pieces together.

97
00:06:18,250 --> 00:06:23,130
The TF-IDF score of feature i for

98
00:06:23,130 --> 00:06:28,360
document j is obtained by multiplying
the Term Frequency and the IDF.

99
00:06:31,600 --> 00:06:36,300
So given a document you
compute the TF-IDF scores.

100
00:06:36,300 --> 00:06:38,450
For every term in the document.

101
00:06:38,450 --> 00:06:43,030
And then you sort all the terms in
the document by their TF-IDF scores.

102
00:06:43,030 --> 00:06:46,750
And then you have some
kind of threshold or

103
00:06:46,750 --> 00:06:50,010
you might pick the set of words
with the highest TF-IDF scores in

104
00:06:50,010 --> 00:06:54,010
the document together with their scores
and that would be the top profile.

105
00:06:54,010 --> 00:06:54,660
So in this case,

106
00:06:54,660 --> 00:06:58,440
a doc profile is a real value vector
as opposed to a boolean vector.

107
00:06:59,870 --> 00:07:04,000
Now that we have item profiles, our
next task is to construct user profiles.

108
00:07:06,510 --> 00:07:10,850
Let's say we have a user who has rated
items with profiles i1 through i n.

109
00:07:13,830 --> 00:07:23,870
Now remember,
I want to i-n our vectors of, of entries.

110
00:07:23,870 --> 00:07:27,890
Let's say this is i-1,
this i-2, i-3 and so on.

111
00:07:29,250 --> 00:07:30,214
And here is i-n.

112
00:07:31,380 --> 00:07:35,280
These are each is a vector
in a high dimensional space,

113
00:07:35,280 --> 00:07:40,600
with many many Now the simplest
way to construct a user

114
00:07:40,600 --> 00:07:46,204
profile from a set of item profiles
is just to average the item profiles.

115
00:07:46,204 --> 00:07:51,172
[SOUND] Where N is the total
number of item profiles.

116
00:07:51,172 --> 00:07:55,766
So if I take all the item profiles
in the users you know, of,

117
00:07:55,766 --> 00:08:00,286
of all the item the user has has rated and
then take the average,

118
00:08:00,286 --> 00:08:04,240
that would be the simplest way
of constructing a user profile.

119
00:08:05,320 --> 00:08:08,550
Now this doesn't take into account
that the user liked certain items

120
00:08:08,550 --> 00:08:09,740
more than others.

121
00:08:09,740 --> 00:08:12,970
So in that case we might want
to use a weighted average,

122
00:08:12,970 --> 00:08:18,530
where the weight is equal to the rating
given by the user for for each item.

123
00:08:18,530 --> 00:08:21,120
Then you'd have a weighted
average item profile.

124
00:08:26,140 --> 00:08:28,990
A variant of this is to
normalize these weights

125
00:08:28,990 --> 00:08:30,890
using the average rating of the user.

126
00:08:30,890 --> 00:08:33,550
And you've seen example
that makes this idea clear.

127
00:08:35,680 --> 00:08:38,750
And of course, much more sophisticated
aggregations are possible.

128
00:08:38,750 --> 00:08:41,089
Here we're only looking at
some very simple examples.

129
00:08:42,680 --> 00:08:45,360
Let's look at an example that you know?

130
00:08:45,360 --> 00:08:48,510
That'll clarify weighted
average item profiles.

131
00:08:48,510 --> 00:08:49,790
And how to normalize weights.

132
00:08:52,760 --> 00:08:55,220
Let's start with an example
of a Boolean Utility Matrix.

133
00:08:56,300 --> 00:08:57,870
What's a Boolean Utility Matrix?

134
00:08:57,870 --> 00:09:02,930
All we have is information of whether
the user purchased an item or not.

135
00:09:02,930 --> 00:09:03,940
For example.

136
00:09:03,940 --> 00:09:06,240
So each entry is either a zero or a one.

137
00:09:08,950 --> 00:09:11,870
Let's say the items are movies and
the only feature is actor.

138
00:09:14,220 --> 00:09:19,110
The item profile in this case is
a vector for zero or one for each actor.

139
00:09:19,110 --> 00:09:21,780
Zero if that actor did
not appear in that movie.

140
00:09:21,780 --> 00:09:23,420
And one if that actor
appeared in that movie.

141
00:09:26,510 --> 00:09:31,630
Suppose user x has watched five movies and
two of those movies feature actor a and

142
00:09:31,630 --> 00:09:33,290
three of those movies feature actor b.

143
00:09:37,310 --> 00:09:40,559
Now the simplest user profile is
just the mean of the item profiles.

144
00:09:42,470 --> 00:09:50,350
Remember there are 5 vectors, and
2 of those have a 1 for feature A.

145
00:09:50,350 --> 00:09:54,480
And so the data feature A is going to
be 2 divided by the total number of

146
00:09:54,480 --> 00:09:57,455
item profiles, which is 5, which is 0.4.

147
00:09:57,455 --> 00:10:01,184
And the weight of feature B,
correspondingly, is going to be 3/5.

148
00:10:03,850 --> 00:10:06,650
Let's look at the more complex
example of its star ratings.

149
00:10:09,510 --> 00:10:12,810
Suppose we have star ratings
in the range of one to five.

150
00:10:12,810 --> 00:10:15,740
And the user has once
again watched five movies.

151
00:10:15,740 --> 00:10:19,300
And there are two movies starring actor
A and three movies starring actor B.

152
00:10:20,580 --> 00:10:24,070
The movies that actor A starred in,
the user rated three and five.

153
00:10:24,070 --> 00:10:29,960
But with the movie that that actor B acted
in, the user rated one, two, and four.

154
00:10:32,620 --> 00:10:37,590
So since we have five star ratings and
the user gives lower ratings for

155
00:10:37,590 --> 00:10:40,630
movies they didn't like and
higher ratings for movies they liked.

156
00:10:40,630 --> 00:10:45,880
It's somewhat apparent from these
ratings that the user liked

157
00:10:45,880 --> 00:10:51,410
at least one of the movies from from Actor
A and one of the movies from Actor B.

158
00:10:51,410 --> 00:10:52,080
But didn't he, but

159
00:10:52,080 --> 00:10:56,870
they really didn't like two of Actor B's
movies, the ones that were rated 1 and 2.

160
00:10:56,870 --> 00:11:00,290
1 and 2 are in fact, negative ratings.

161
00:11:00,290 --> 00:11:01,810
Not positive ratings.

162
00:11:01,810 --> 00:11:03,850
And we try to capture this fact.

163
00:11:03,850 --> 00:11:07,240
The idea of normalizing ratings
helps us capture the idea that

164
00:11:07,240 --> 00:11:10,880
some ratings are actually negative
ratings and some are positive ratings.

165
00:11:10,880 --> 00:11:14,060
But the baseline, you know,
users are very different from each other.

166
00:11:14,060 --> 00:11:16,820
Some users are just more generous
in their ratings than others.

167
00:11:16,820 --> 00:11:19,750
So, for, user a, for instance.

168
00:11:19,750 --> 00:11:22,850
A four might be a widely positive rating.

169
00:11:22,850 --> 00:11:26,780
But if for another,
four might just be an average rating.

170
00:11:26,780 --> 00:11:29,340
To sort of capture this idea,

171
00:11:29,340 --> 00:11:33,750
we want to baseline each user's
ratings by their average rating.

172
00:11:33,750 --> 00:11:37,600
So in this case, the,
this user's average rating is a three.

173
00:11:37,600 --> 00:11:41,810
If you, average all the five ratings
that the user, has provided,

174
00:11:41,810 --> 00:11:44,420
the average rating, is a three.

175
00:11:44,420 --> 00:11:48,230
And so what we're going to do is just
subtract the average rating from each of

176
00:11:48,230 --> 00:11:51,150
the individual movie ratings.

177
00:11:51,150 --> 00:11:55,520
So in this case the movies with
actor A the normalized ratings in

178
00:11:55,520 --> 00:11:59,860
that case a three and five,
become zero and plus two.

179
00:11:59,860 --> 00:12:07,320
And for actor B, the normalized ratings
become minus two, minus one, and plus one.

180
00:12:07,320 --> 00:12:12,230
Notice that this captures intuition
that the user did not like, the,

181
00:12:12,230 --> 00:12:16,760
the first two movies with actor B,
whereas he really liked, the, the,

182
00:12:16,760 --> 00:12:19,170
the second movie with, with, with actor A.

183
00:12:19,170 --> 00:12:23,370
Where the first movie with actor A was,
you know, was kind of an average movie.

184
00:12:25,490 --> 00:12:28,310
Once you do this normalization,
then you can compute the profile,

185
00:12:28,310 --> 00:12:29,950
the profile weights.

186
00:12:29,950 --> 00:12:34,060
But in this case we divide not
by the total number of movies.

187
00:12:34,060 --> 00:12:38,510
But by the total number of
movies with a specific feature.

188
00:12:38,510 --> 00:12:41,990
So in this case there
are two movies with actor A.

189
00:12:41,990 --> 00:12:43,890
And profile weight for

190
00:12:43,890 --> 00:12:50,260
actor A the feature with actor A is zero
plus two divided by two which is one.

191
00:12:50,260 --> 00:12:55,730
And similarly the feature actor
B has a profile weight of -2/3.

192
00:12:55,730 --> 00:13:00,420
This indicates a mild positive
preference for, for actor A.

193
00:13:00,420 --> 00:13:02,720
And a mild negative preference for
actor B.

194
00:13:04,090 --> 00:13:05,430
Now that we have user profiles and

195
00:13:05,430 --> 00:13:09,140
actor profiles, the next task is to
recommend certain items to the user.

196
00:13:10,580 --> 00:13:16,250
The key step in this is to take a pair
of user profile and item profile, and

197
00:13:16,250 --> 00:13:21,210
figure out what the rating for
that user and item pair is likely to be.

198
00:13:23,220 --> 00:13:25,390
Remember that both the user profile and

199
00:13:25,390 --> 00:13:28,740
the item profile are vectors
in high-dimensional space.

200
00:13:28,740 --> 00:13:32,040
In this case I've shown them in a
two-dimensional space, when the reality of

201
00:13:32,040 --> 00:13:34,010
course, they're embedded in
a much higher dimensional space.

202
00:13:35,020 --> 00:13:38,406
You might recall from a prior lecture
that when you have vectors in

203
00:13:38,406 --> 00:13:42,471
higher dimensional space a good distance
metric between the pair of vectors is

204
00:13:42,471 --> 00:13:44,825
the angle theta between
the pair of vectors.

205
00:13:46,927 --> 00:13:52,020
In particular, you can estimate
the angle using the cosine formula.

206
00:13:52,020 --> 00:13:55,890
The cosine of Theta, Theta, the angle
between the two vectors is given by

207
00:13:55,890 --> 00:13:59,640
the dot product of the two vectors,
divided by the product of the magnitudes.

208
00:14:01,290 --> 00:14:04,280
And this distance, in, in this case,

209
00:14:04,280 --> 00:14:08,810
we'll call this cosine similarity between,
the user x and the item i.

210
00:14:10,400 --> 00:14:14,020
Now technically the cosine distance
is actually the angle theta and

211
00:14:14,020 --> 00:14:15,250
not the cosine of the angle.

212
00:14:16,370 --> 00:14:16,920
Right?

213
00:14:16,920 --> 00:14:21,660
The cosine distance, as we studied in
an earlier lecture, is the angle theta and

214
00:14:21,660 --> 00:14:25,130
the cosine similarity is
the angle 180 minus theta.

215
00:14:26,580 --> 00:14:30,350
Now the smaller the angle,
the more similar the item x and

216
00:14:30,350 --> 00:14:36,630
the the, the more similar the user x and
the item i r.

217
00:14:36,630 --> 00:14:42,750
And therefore the similarity 180
minus data is going to be larger.

218
00:14:45,954 --> 00:14:50,513
But for convenience, were going to
actually use the cosine of theta as,

219
00:14:50,513 --> 00:14:52,610
as a similarity measure.

220
00:14:52,610 --> 00:14:57,590
Notice that as the angle of theta becomes
smaller, cost theta becomes larger.

221
00:14:57,590 --> 00:15:02,430
And as it angle theta becomes larger and
larger, the cosine becomes smaller and

222
00:15:02,430 --> 00:15:05,540
smaller in fact,
as theta becomes greater than 90,

223
00:15:05,540 --> 00:15:07,049
the cosine of theta becomes negative.

224
00:15:08,100 --> 00:15:11,690
And, so this captures intuition, that,
as the angle becomes smaller and

225
00:15:11,690 --> 00:15:15,800
smaller, x and i are more and
more similar to each other, and the, and

226
00:15:15,800 --> 00:15:18,920
it's more likely that x will
give a high rating to item i.

227
00:15:22,490 --> 00:15:24,089
So the way we make
predictions is as follows.

228
00:15:25,280 --> 00:15:30,670
Given the user x, we compute the cosine
similarity between that user and

229
00:15:30,670 --> 00:15:32,690
all the items in the catalog.

230
00:15:32,690 --> 00:15:35,930
And then you pick the items with
the highest cosine similarity and

231
00:15:35,930 --> 00:15:36,980
recommend those to the user.

232
00:15:40,410 --> 00:15:43,729
And that's a theory of
content-based recommendations.

233
00:15:44,730 --> 00:15:46,010
Now let's look at some of the pros and

234
00:15:46,010 --> 00:15:52,810
cons of the content-based
recommendation approach.

235
00:15:52,810 --> 00:15:56,120
The biggest pro of the content-based
recommendation approach is that you don't

236
00:15:56,120 --> 00:16:00,570
need data about other users in order to
make recommendations to a specific user.

237
00:16:01,940 --> 00:16:05,059
This turns out to be a very,
very good thing because you know,

238
00:16:05,059 --> 00:16:09,686
you can start working making content-based
recommendations from day one for, for

239
00:16:09,686 --> 00:16:10,771
your very first user.

240
00:16:15,643 --> 00:16:19,267
Another good thing about content-based
recommendation is that it can recommend to

241
00:16:19,267 --> 00:16:20,449
users a very unique taste.

242
00:16:21,780 --> 00:16:23,950
When we go,
when we get to collaborative filtering.

243
00:16:23,950 --> 00:16:28,360
We'll see that collaborative filtering
can make recommendations to a user.

244
00:16:28,360 --> 00:16:30,060
We need to find similar users.

245
00:16:31,680 --> 00:16:34,470
The problem with that is if
the user were very unique or

246
00:16:34,470 --> 00:16:37,328
idiosyncratic taste there may
not be any other similar users.

247
00:16:37,328 --> 00:16:41,410
But the content-based approach is
able to deal marginally with this,

248
00:16:41,410 --> 00:16:42,820
with the fact that it can make.

249
00:16:42,820 --> 00:16:46,010
You know, user can very
unique tastes as long as the,

250
00:16:46,010 --> 00:16:49,350
we can build item profiles for
the items that the user likes.

251
00:16:49,350 --> 00:16:50,290
And a user profile for

252
00:16:50,290 --> 00:16:53,020
the user based on that,
we can make recommendations to that user.

253
00:16:54,150 --> 00:16:58,100
The third row is that we're able to
recommend new and unpopular items.

254
00:16:58,100 --> 00:17:02,350
Now when a new item comes in we
don't need any ratings from users to

255
00:17:02,350 --> 00:17:03,930
build the item profile.

256
00:17:03,930 --> 00:17:07,390
The item profile depends entirely
on the features of the items and

257
00:17:07,390 --> 00:17:11,365
not on how other users rated the item so
we don't have a so called

258
00:17:11,365 --> 00:17:16,090
first-rater problem that we'll see in the
in the collaborative filtering approach.

259
00:17:16,090 --> 00:17:19,159
We can make recommendation for
an item as soon as it becomes available.

260
00:17:21,980 --> 00:17:25,460
And finally, whenever the content-based
approach makes a recommendation,

261
00:17:25,460 --> 00:17:30,470
you can provide an explanation to the user
for why a certain item was recommended.

262
00:17:30,470 --> 00:17:33,610
In particular, you can just the list the
content feature that caused the item to

263
00:17:33,610 --> 00:17:34,600
be recommended.

264
00:17:34,600 --> 00:17:39,400
For example, if you,
recommend a news article to a user, for

265
00:17:39,400 --> 00:17:41,520
example, using a content-based approach.

266
00:17:41,520 --> 00:17:45,050
You might be able to say look in
the past you've spent a lot of

267
00:17:45,050 --> 00:17:49,230
time reading articles that mention Syria.

268
00:17:49,230 --> 00:17:52,060
And that's why I'm recommending
this article on Syria to you.

269
00:17:56,000 --> 00:17:58,350
So these have some of the pros
of the content-based approach.

270
00:17:58,350 --> 00:17:59,420
But now let's look at the cons.

271
00:18:01,840 --> 00:18:04,910
The most important problem or

272
00:18:04,910 --> 00:18:07,440
the most serious problem with
a content based approach.

273
00:18:07,440 --> 00:18:11,000
Is that finding the appropriate
features is very, very hard.

274
00:18:11,000 --> 00:18:14,749
For example how do you find features for
images.

275
00:18:17,030 --> 00:18:18,830
Or movies, or music.

276
00:18:18,830 --> 00:18:23,340
Now in the case of movies we suggested
a set of features that include actors and

277
00:18:23,340 --> 00:18:24,946
directors and so on.

278
00:18:24,946 --> 00:18:30,396
But it turns out that movies often
[INAUDIBLE] genres and users are n.ot very

279
00:18:30,396 --> 00:18:36,106
often loyal to specific actors or
directors and the similar

280
00:18:36,106 --> 00:18:42,076
case of music it's very hard to you
know boxed music in specific genres and

281
00:18:42,076 --> 00:18:47,640
musicians and so on, and
of course, are very hard to find.

282
00:18:47,640 --> 00:18:54,421
So in general, the finding of features
to make content-based a very very hard

283
00:18:54,421 --> 00:19:00,200
problem, reason why the content-based
approach is not more popular.

284
00:19:03,510 --> 00:19:06,190
The second problem is one
of overspecialization.

285
00:19:06,190 --> 00:19:10,910
Remember, the user profile is built using,
the item profile of the,

286
00:19:10,910 --> 00:19:13,090
the items that the user has rated or
purchased.

287
00:19:15,360 --> 00:19:19,370
Now, because of this, if a user has
never rated a certain kind of movie or

288
00:19:19,370 --> 00:19:21,440
a certain genre of movie.

289
00:19:21,440 --> 00:19:25,590
He will never be recommended a movie
in that, in that genre for example.

290
00:19:25,590 --> 00:19:28,780
Or he'll never be recommended
a piece of music,

291
00:19:28,780 --> 00:19:31,510
that's outside his previous, preferences.

292
00:19:32,660 --> 00:19:34,870
In general,
people might have multiple interests, and

293
00:19:34,870 --> 00:19:37,410
might express only some
of them in the past.

294
00:19:37,410 --> 00:19:43,960
And so, it's hard to, you know, so it'd
be, Easy this way to miss recommending

295
00:19:43,960 --> 00:19:49,920
interesting items users because you don't
have enough, eh, enough user on the user.

296
00:19:49,920 --> 00:19:52,670
Another serious problem of
the content-based approach is

297
00:19:52,670 --> 00:19:56,350
that it's unable to exploit
the quality judgments of other users.

298
00:19:56,350 --> 00:19:58,600
For example there might
be a certain video or

299
00:19:58,600 --> 00:20:04,490
movie that's widely popular Across a,
you know, wide cross-section of users.

300
00:20:04,490 --> 00:20:06,840
.
However, the current user has not

301
00:20:06,840 --> 00:20:09,070
expressed interest in that kind of movie.

302
00:20:09,070 --> 00:20:10,160
And therefore the content they

303
00:20:10,160 --> 00:20:12,360
support should never recommend
that movie to that user.

304
00:20:16,470 --> 00:20:19,290
A final problem that we have with
a content-based approach is one of a, a,

305
00:20:19,290 --> 00:20:20,869
a cold-start problem for new users.

306
00:20:22,090 --> 00:20:25,780
Remember, the user profile
is built by aggregating item

307
00:20:25,780 --> 00:20:29,150
profiles of the items the user has rated.

308
00:20:29,150 --> 00:20:32,310
When you have a new user,
the new user has not rated any items, and

309
00:20:32,310 --> 00:20:36,060
so the, so, so there is no user profile.

310
00:20:36,060 --> 00:20:39,669
So, there is a challenging problem of how
to build a user profile for a new user.

311
00:20:41,530 --> 00:20:47,090
In most practical situations new
users start with you know most

312
00:20:47,090 --> 00:20:52,290
recommended systems start off new users
with some kind of average profile based on

313
00:20:52,290 --> 00:20:57,450
a system wide average And then over time
user profile evolves as rates more and

314
00:20:57,450 --> 00:21:00,150
more items and
becomes more individualized to the user

