1
00:00:00,320 --> 00:00:04,540
Anant has already talked to you about
recommender systems, and so far we have

2
00:00:04,540 --> 00:00:08,590
learned how to do, collaborative
filtering recommender systems and

3
00:00:08,590 --> 00:00:11,900
how to do,
profile based recommender systems.

4
00:00:11,900 --> 00:00:14,090
What we will do today in this lecture,

5
00:00:14,090 --> 00:00:17,280
we will look into latent
factor recommender models.

6
00:00:17,280 --> 00:00:20,090
So basically the idea will
be that we will adopt this

7
00:00:20,090 --> 00:00:24,550
machine learning perspective where we'll
be using optimization in order to build.

8
00:00:24,550 --> 00:00:27,080
Very accurate recommender systems.

9
00:00:27,080 --> 00:00:31,490
And in particular, our lecture
today will be moderated by, by,

10
00:00:31,490 --> 00:00:33,570
by the Netflix price, right?

11
00:00:33,570 --> 00:00:38,990
So, Netflix is, is a company that
is renting out, DVDs, and in order

12
00:00:38,990 --> 00:00:43,770
to increase customer experience, Netflix
wants to be able to recommend which

13
00:00:43,770 --> 00:00:47,970
customers will like which movies, and this
is called the movie recommender system.

14
00:00:47,970 --> 00:00:51,580
Right, and a few years ago, actually,
Netflix prepared some data and

15
00:00:51,580 --> 00:00:55,890
created The Grand Challenge,
where the idea was who can build the best

16
00:00:55,890 --> 00:01:00,550
recommender system and the person who does
that gets a one million dollars prize.

17
00:01:00,550 --> 00:01:01,590
So here is how the.

18
00:01:02,870 --> 00:01:04,340
Competition was structured.

19
00:01:04,340 --> 00:01:08,570
So the idea was that first we would get,
scientists would get training data.

20
00:01:08,570 --> 00:01:12,740
This is a set of data of 100 million
ratings from 480,000, users.

21
00:01:12,740 --> 00:01:15,430
And around 17,000 movies.

22
00:01:15,430 --> 00:01:18,303
And this is six years' worth of data.

23
00:01:18,303 --> 00:01:19,700
From 2000 to 2005.

24
00:01:19,700 --> 00:01:20,240
Right?

25
00:01:20,240 --> 00:01:24,620
So based on this data
the idea is that we build.

26
00:01:24,620 --> 00:01:28,290
Or create a recommender system and
then we would send these

27
00:01:28,290 --> 00:01:33,040
recommendations to Netflix that would
check our predicted recommendation.

28
00:01:33,040 --> 00:01:35,960
Basically how much does
a given user like the movie.

29
00:01:35,960 --> 00:01:42,030
We did their test data set which is a set
of Ratings that only Netflix has, and

30
00:01:42,030 --> 00:01:47,670
this way we would want to predict or, the,
and make sure that the predicted ratings

31
00:01:47,670 --> 00:01:52,340
are correspond to the actual true ratings
of the people who like those movies.

32
00:01:52,340 --> 00:01:53,660
So that's basically the idea.

33
00:01:53,660 --> 00:01:56,140
So, we are given hundred
million movie ratings, and

34
00:01:56,140 --> 00:01:59,370
we want to build a recommender
system out of this.

35
00:01:59,370 --> 00:02:02,890
So, the way we can think
of our ratings data in,

36
00:02:02,890 --> 00:02:05,220
is in terms of the utility matrix R.

37
00:02:05,220 --> 00:02:08,730
And, basically,
the way this matrix is composed is that

38
00:02:08,730 --> 00:02:12,540
rows corresponds to movies and
columns corresponds to users.

39
00:02:12,540 --> 00:02:13,460
And this matrix,

40
00:02:13,460 --> 00:02:18,520
most of it is empty in a sense that most
users haven't seen most of the movies.

41
00:02:18,520 --> 00:02:21,610
Right?
But whenever a mo, a user sees a movie and

42
00:02:21,610 --> 00:02:25,360
rates that movie, that represents
an entry in this matrix, right?

43
00:02:25,360 --> 00:02:29,730
So for example, this entry,
in this particular en,

44
00:02:29,730 --> 00:02:34,880
element of the matrix, number 5,
that would mean that, a user in the,

45
00:02:34,880 --> 00:02:40,780
in the fourth column liked the third movie
a lot and rated it given five stars.

46
00:02:40,780 --> 00:02:42,380
Right?
So now given such a,

47
00:02:42,380 --> 00:02:45,420
such a matrix, we want to formulate.

48
00:02:45,420 --> 00:02:46,970
The recommendation problem.

49
00:02:46,970 --> 00:02:50,410
The way we think of the recommendation
problem is in a sense that we want to

50
00:02:50,410 --> 00:02:54,890
predict how much will a given
user like some movie that

51
00:02:54,890 --> 00:02:56,320
they will see in the future.

52
00:02:56,320 --> 00:03:01,640
So in this sense, we can use the
historical rating data to build our model.

53
00:03:01,640 --> 00:03:04,090
This is what we'll call the training data.

54
00:03:04,090 --> 00:03:09,330
And then using this strange model we
want to predict how much the users like

55
00:03:09,330 --> 00:03:13,000
movies that they haven't yet seen and
we will call this the test data.

56
00:03:13,000 --> 00:03:17,330
And now when I say that we want to predict
how much will a user like a given movie

57
00:03:17,330 --> 00:03:22,450
the way we want to quantify how accurate
is our prediction we will use what is

58
00:03:22,450 --> 00:03:24,600
called the root mean square data.

59
00:03:24,600 --> 00:03:29,040
So, basically the idea will be that our
goal is to make predictions about how

60
00:03:29,040 --> 00:03:34,170
much users are going to like the movies,
then we will go and ask these users hey,

61
00:03:34,170 --> 00:03:36,730
how much do you really like that movie?

62
00:03:36,730 --> 00:03:38,710
And then we want these two,
the prediction and

63
00:03:38,710 --> 00:03:41,210
the true value,
to be as close as possible.

64
00:03:41,210 --> 00:03:45,230
So, the root mean square error,
the way the formula works, is follows.

65
00:03:45,230 --> 00:03:45,990
Right?

66
00:03:45,990 --> 00:03:52,495
For a given user x in movie i I
want to predict how much will that

67
00:03:52,495 --> 00:03:57,500
user-movie pair what will be
the rating of user x to movie i.

68
00:03:57,500 --> 00:04:02,610
Of course I also know the true rating for
movie i of user x.

69
00:04:02,610 --> 00:04:06,450
And what we want to do is we want the
discrepancy between the predicted rating

70
00:04:06,450 --> 00:04:07,560
and the true rating.

71
00:04:07,560 --> 00:04:10,680
To be, to be small in a sense that
we want to take the difference

72
00:04:11,760 --> 00:04:13,150
square this difference.

73
00:04:13,150 --> 00:04:13,770
The reason we want to

74
00:04:13,770 --> 00:04:16,450
square is because the difference
could be either positive of negative.

75
00:04:16,450 --> 00:04:19,360
So if we square it up,
everything is positive.

76
00:04:19,360 --> 00:04:23,780
Now we sum all these errors across
all the users and all the movies.

77
00:04:23,780 --> 00:04:27,280
We take the square root and
divide by the total number of ratings, and

78
00:04:27,280 --> 00:04:29,640
this is called root-mean-square error.

79
00:04:29,640 --> 00:04:35,930
So in some sense, our goal is to in
some sense complete this matrix R.

80
00:04:35,930 --> 00:04:39,780
Wherever the values for
ratings are unknown.

81
00:04:39,780 --> 00:04:42,980
And our goal is to complete this
matrix as accurately as possible.

82
00:04:42,980 --> 00:04:43,930
So in some sense,

83
00:04:43,930 --> 00:04:48,990
we want to predict how much is a given
user going to like a given movie.

84
00:04:50,340 --> 00:04:53,300
So what is the input and
what is the output, right?

85
00:04:53,300 --> 00:04:58,140
So what Netflix gives us is the training
data of hun, 100 million ratings.

86
00:04:58,140 --> 00:05:02,970
So half a million people,
20,000 movies, and, and ratings.

87
00:05:02,970 --> 00:05:07,140
Our goal is to predict the,
a few ratings for every user.

88
00:05:07,140 --> 00:05:10,680
Where the evaluation criteria or
the good, the way, how we

89
00:05:10,680 --> 00:05:15,230
evaluate how good that our predictions,
are based on the root mean squared error.

90
00:05:15,230 --> 00:05:20,240
And, what we also know, for example, is
that, that the root mean square error of

91
00:05:20,240 --> 00:05:25,390
the Netflix system that was used in house
and at that point in time was zero, 0.95.

92
00:05:25,390 --> 00:05:26,830
Okay?

93
00:05:26,830 --> 00:05:31,155
So in some sense,
on average this, this in,

94
00:05:31,155 --> 00:05:36,940
in-house recommender system was about one
star away from the, from the true rating.

95
00:05:36,940 --> 00:05:37,610
Right?

96
00:05:37,610 --> 00:05:42,178
So, this competition that was
taking place in 2009 and 2010.

97
00:05:42,178 --> 00:05:46,140
More than 2,700 teams
joined the competition.

98
00:05:46,140 --> 00:05:48,950
And what matrix also said
was whoever is able to

99
00:05:48,950 --> 00:05:54,480
improve their in-house recommended
system by 10% gets $1 million prize.

100
00:05:54,480 --> 00:05:57,220
Right?
So the idea is how can

101
00:05:57,220 --> 00:06:00,060
we beat the in-house
recommended system and get.

102
00:06:00,060 --> 00:06:03,289
.
The root brings square error from .9 to

103
00:06:03,289 --> 00:06:05,580
.85 that's the whole idea.

104
00:06:06,580 --> 00:06:10,820
Now the structure of how to
do this is the following.

105
00:06:10,820 --> 00:06:14,600
Basically a modern recommender
system today is composed of

106
00:06:14,600 --> 00:06:16,650
several different components.

107
00:06:16,650 --> 00:06:19,895
And, the idea is that basically
we want to adopt this kind of

108
00:06:19,895 --> 00:06:23,880
multi-scale modeling approach to the,
to the rating data.

109
00:06:23,880 --> 00:06:28,200
And the idea is that we want to combine
several different views of the data from

110
00:06:28,200 --> 00:06:33,880
this type of global top level view of the
data to some kind of regional modeling of

111
00:06:33,880 --> 00:06:39,300
the data, all the way down to very local
interactions between movies and the users.

112
00:06:39,300 --> 00:06:40,760
Right?
And what the,

113
00:06:40,760 --> 00:06:42,580
the goal of today's lecture will be,

114
00:06:42,580 --> 00:06:46,710
that we basically focus on the,
on the, original or the,

115
00:06:46,710 --> 00:06:51,480
you know, the middle level, modeling of
the data through matrix factorization.

116
00:06:51,480 --> 00:06:56,550
So the methodology that we'll be using to,
to achieve, our modeling will be

117
00:06:56,550 --> 00:07:00,540
through matrix factorization which relates
very nicely to the singular value,

118
00:07:00,540 --> 00:07:02,500
the composition and
dimensionality reduction.

119
00:07:03,950 --> 00:07:06,500
Lecture that we talked about previously.

120
00:07:06,500 --> 00:07:10,700
But before we kind of go to the main body
of today's lecture, let me talk to you

121
00:07:10,700 --> 00:07:15,430
about global modeling and also about
collaborative filtering which is very

122
00:07:15,430 --> 00:07:20,090
a local way of modeling the data So what
do I mean by global modeling of the data.

123
00:07:21,220 --> 00:07:24,270
The global modeling of the data is
basically that you want to kind of

124
00:07:24,270 --> 00:07:29,510
compute some average, average rating over
the movies and, and also over the users.

125
00:07:29,510 --> 00:07:33,030
Right so if you would use a global
way to model how much will

126
00:07:33,030 --> 00:07:38,200
a given user like a given movie then we
could say the min movie rating in the,

127
00:07:38,200 --> 00:07:41,571
in our data set on Netflix is 3.7 stars.

128
00:07:41,571 --> 00:07:45,390
And then we could say oh but for
example the Sixth Sense movie.

129
00:07:45,390 --> 00:07:49,090
Is, is rated half a star
above the average rating.

130
00:07:49,090 --> 00:07:52,460
And then we could say, oh, but
this given user is very critical.

131
00:07:52,460 --> 00:07:55,040
So they tend to, they tend to rate things

132
00:07:56,250 --> 00:08:00,380
about 0.2 stars below the average rating,
right.

133
00:08:00,380 --> 00:08:04,880
So we could now say, okay,
so 3.7 is our baseline.

134
00:08:04,880 --> 00:08:10,010
Now we have a movie that is above average
for half a star, so that is 4.2, but this

135
00:08:10,010 --> 00:08:16,010
person, Joe, is a critical person, so they
tend to be .2 stars below the average.

136
00:08:16,010 --> 00:08:19,890
So, 4.2 minus, minus .2 equals 4, right?

137
00:08:19,890 --> 00:08:24,840
So, in some sense our prediction how
much will Joe like the Sixth Sense,

138
00:08:24,840 --> 00:08:30,895
Sense would be 4 stars,
as I said because of 3.7 plus .5.

139
00:08:30,895 --> 00:08:33,910
Minus 0.2 equals 4.0, right?

140
00:08:33,910 --> 00:08:38,880
So this is how basically,
just based on the average,

141
00:08:38,880 --> 00:08:40,850
average of the user rating and

142
00:08:40,850 --> 00:08:45,050
average of the movie rating and overall
average rating throughout the data set,

143
00:08:45,050 --> 00:08:49,720
we can already make a prediction how
much will a user like a given movie.

144
00:08:49,720 --> 00:08:52,150
Of course this is a very crude,
very global prediction.

145
00:08:52,150 --> 00:08:54,530
But it turns out to be very useful.

146
00:08:54,530 --> 00:08:57,280
So this is at the end,
one end of the spectrum.

147
00:08:57,280 --> 00:09:01,870
On the other end of the spectrum we can
use very kind of local neighborhood

148
00:09:01,870 --> 00:09:06,220
properties of a given user and
a given movie to make recommendations.

149
00:09:06,220 --> 00:09:08,440
And this is what
collaborative filtering does.

150
00:09:08,440 --> 00:09:11,080
Right?
So in some sense the idea is,

151
00:09:11,080 --> 00:09:12,480
would be following, right?

152
00:09:12,480 --> 00:09:13,830
So I would like to,

153
00:09:13,830 --> 00:09:18,740
to estimate what are other movies that
are similar to the Sixth Sense that Joe

154
00:09:18,740 --> 00:09:23,450
also rated and based on this I
would then make a recommendation.

155
00:09:23,450 --> 00:09:23,980
Right?

156
00:09:23,980 --> 00:09:28,380
And based on these kinds of signals I
would then make a final estimate that Joe

157
00:09:28,380 --> 00:09:33,620
is like, going to like our movie with,
or rate it with 3.8 stars.

158
00:09:33,620 --> 00:09:34,250
Okay?

159
00:09:34,250 --> 00:09:37,760
So how does collaborative
filtering do what I just said?

160
00:09:37,760 --> 00:09:39,400
The idea is basically the following.

161
00:09:39,400 --> 00:09:42,830
Is that we want to derive
unknown ratings from those,

162
00:09:42,830 --> 00:09:44,680
from the ratings of similar movies.

163
00:09:44,680 --> 00:09:46,280
Right?
This is the item-item base

164
00:09:46,280 --> 00:09:47,760
collaborative filtering.

165
00:09:47,760 --> 00:09:51,730
So now the, the big question
with collaborative filtering.

166
00:09:51,730 --> 00:09:55,370
Is how do we operationalize
the notion of similar?

167
00:09:55,370 --> 00:09:57,050
Right?
So, I basically won't need

168
00:09:57,050 --> 00:10:02,490
to define a similarity measure that
tells me how similar are movies i and j.

169
00:10:02,490 --> 00:10:04,830
And now that I, that for
every pair of movies for

170
00:10:04,830 --> 00:10:07,140
example, I'm able to
measure the similarity.

171
00:10:07,140 --> 00:10:11,900
Then the idea is the following, I given
that I want to make up a recommendation.

172
00:10:11,900 --> 00:10:16,980
For movie I, the way I will do this is
I would go find K nearest neighbors,

173
00:10:16,980 --> 00:10:21,280
basically K most similar movies
to movie I, and then I would ask,

174
00:10:21,280 --> 00:10:26,520
ask the ratings and, and add them together
in a weighted combination where movies

175
00:10:26,520 --> 00:10:29,910
that are more similar to my movie I,
they have a higher weight.

176
00:10:29,910 --> 00:10:32,630
So, this is the formula that I have.

177
00:10:32,630 --> 00:10:34,540
At the bottom, So let me explain it.

178
00:10:34,540 --> 00:10:35,070
Right.
So

179
00:10:35,070 --> 00:10:39,630
what we want to do is we want to predict
how much will user x like movie i.

180
00:10:39,630 --> 00:10:41,790
The way we do this is that for a.

181
00:10:41,790 --> 00:10:43,960
For a given movie.

182
00:10:43,960 --> 00:10:48,980
And for we go and
find let's say k-nearest neighbors.

183
00:10:48,980 --> 00:10:53,360
Meaning k most similar movies
j to this query movie i.

184
00:10:53,360 --> 00:10:59,770
And then we say what was the rating that
this user X gave to that given movie I,

185
00:10:59,770 --> 00:11:04,120
and then we also multiply this with
the similarity to movie I and J right.

186
00:11:04,120 --> 00:11:09,780
So the idea basically is that if movies I
and J are very similar, then SIJ will be

187
00:11:09,780 --> 00:11:15,620
will be high so the rating that user X
gave to movie J will have a high weight.

188
00:11:15,620 --> 00:11:20,280
If I, if the movie J is of close
similarity to I, then what will happen is

189
00:11:20,280 --> 00:11:27,170
that, that rating of user X to the movie
J will have no weight in this summation.

190
00:11:27,170 --> 00:11:30,780
And then, at the bottom,
we just normalize, right?

191
00:11:30,780 --> 00:11:33,900
And this is one way how we can compute,

192
00:11:33,900 --> 00:11:36,800
the rating based on
the collaborative filtering.

193
00:11:36,800 --> 00:11:41,440
Now how do we put together
the local effects and

194
00:11:41,440 --> 00:11:45,420
the global effects right how do we
combine collaborative filtering with glo,

195
00:11:45,420 --> 00:11:50,010
with the global effects model that I
described described two slides ago.

196
00:11:50,010 --> 00:11:51,630
The way we do this is.

197
00:11:51,630 --> 00:11:53,870
That basically we just
add the two together.

198
00:11:53,870 --> 00:11:54,940
Right?
So here's,

199
00:11:54,940 --> 00:11:57,700
here's the formula that
I will explain again.

200
00:11:57,700 --> 00:11:59,740
All right so my recommendation or

201
00:11:59,740 --> 00:12:03,750
prediction of how much will
user i user x like movie i.

202
00:12:03,750 --> 00:12:06,330
Is the, the baseline predictor.

203
00:12:06,330 --> 00:12:08,530
Which is what we talked before.

204
00:12:08,530 --> 00:12:10,880
Right.
So the, the baseline predictor for

205
00:12:10,880 --> 00:12:16,510
user x and movie i is simply the average
movie rating plus the average

206
00:12:16,510 --> 00:12:21,180
deviation for user x plus
the average deviation for movie i.

207
00:12:21,180 --> 00:12:24,750
Right.
So in our previous case v of x was

208
00:12:24,750 --> 00:12:27,439
minus point two because our user,
Joe, was.

209
00:12:28,540 --> 00:12:33,590
Very critical, and in our case
the 6th Sense was a good movie, so

210
00:12:33,590 --> 00:12:38,992
the bias of that movie was .5,
and the overall rating was 3.7.

211
00:12:38,992 --> 00:12:39,936
Right?

212
00:12:39,936 --> 00:12:44,250
So our baseline recommendation for
Joe at that point in time was 4.0.

213
00:12:44,250 --> 00:12:48,530
And now, of course,
this is our baseline recommendation.

214
00:12:48,530 --> 00:12:53,470
And now we can take the part from the
collaborative filtering, where we take k

215
00:12:53,470 --> 00:12:58,460
near, closest movie is,
j, to the given movie, I.

216
00:12:58,460 --> 00:13:02,430
And ask, what is the deviation of the,

217
00:13:02,430 --> 00:13:07,710
of the user's rating with regard to the,
to their baseline for that movie.

218
00:13:07,710 --> 00:13:13,870
And take away the combination of this and
compute the recommendation right so

219
00:13:13,870 --> 00:13:18,580
here the the first part is the is
the global part of the model and

220
00:13:18,580 --> 00:13:22,050
the the second part is the collaborative
filtering part of the model.

221
00:13:23,460 --> 00:13:27,770
Of course here are a few caveats
a few things to to be careful about.

222
00:13:27,770 --> 00:13:32,170
First, first one is that finding coming
up with the good similarity measure.

223
00:13:32,170 --> 00:13:37,370
Basically the good SIJ is very hard and
sometimes it's, it's arbitrary or

224
00:13:37,370 --> 00:13:41,860
one has to try many different
metrics before one starts to work.

225
00:13:43,590 --> 00:13:47,320
But pairwise, the problem with pairwise
similarities is that basically they,

226
00:13:47,320 --> 00:13:51,260
they, they Consider the similarities
between the movies but

227
00:13:51,260 --> 00:13:54,650
they kind of ignore
the similarities between the users.

228
00:13:54,650 --> 00:13:57,710
And another thing here we
could say is that taking this

229
00:13:57,710 --> 00:14:01,300
weighted average sometimes can be,
can be restricted.

230
00:14:01,300 --> 00:14:03,910
But these are kind of,
if there are some issues with this model.

231
00:14:03,910 --> 00:14:05,610
These are the three to talk about.

232
00:14:05,610 --> 00:14:08,360
But in practice this
already works quite well.

233
00:14:08,360 --> 00:14:12,280
So let me show you how I do these
different models work, okay?

234
00:14:12,280 --> 00:14:14,970
So what I'm showing you
with this yellow arrow is,

235
00:14:14,970 --> 00:14:19,200
what is the performance of various
methods on this Netflix price.

236
00:14:19,200 --> 00:14:22,400
Right, for example, if for
a given user we would just go and

237
00:14:22,400 --> 00:14:25,290
predict the global average,
so we would just compute.

238
00:14:25,290 --> 00:14:28,910
what is the global average
star rating on the whole.

239
00:14:28,910 --> 00:14:33,040
Netflix data set and what whenever
a user comes you could just say 3.7.

240
00:14:33,040 --> 00:14:34,140
That's our prediction.

241
00:14:34,140 --> 00:14:38,240
If you do that our root and
square add on would be 1.12.

242
00:14:38,240 --> 00:14:40,390
Right so if we are kind of very naive and

243
00:14:40,390 --> 00:14:43,800
don't even look at the user just
always give the same recommendation.

244
00:14:43,800 --> 00:14:47,120
Our root-mean-square error would be 1.12.

245
00:14:47,120 --> 00:14:51,510
If for example for every user, we would
say we don't even care what is the,

246
00:14:51,510 --> 00:14:54,470
the movie this user is going to, to watch.

247
00:14:54,470 --> 00:14:55,230
We will just go and

248
00:14:55,230 --> 00:15:00,200
predict the average rating for that movie,
that already decreases the error, right.

249
00:15:00,200 --> 00:15:04,450
So, if for every user we just
predict what's their average rating.

250
00:15:04,450 --> 00:15:08,420
That will reduces the error to 1.06.

251
00:15:08,420 --> 00:15:14,240
Now for example told you before that the
Netflix recommender system was at 0.95.

252
00:15:14,240 --> 00:15:20,240
Right so kind of much better than
this naive global, global model.

253
00:15:20,240 --> 00:15:25,130
But, for example, the simple collaborative
filtering that I told you before,

254
00:15:25,130 --> 00:15:27,560
already beats the, the Netflix model.

255
00:15:27,560 --> 00:15:30,790
So the collaborative filtering
that we just talked about,

256
00:15:30,790 --> 00:15:32,390
gives you, the score of 0.94.

257
00:15:32,390 --> 00:15:38,900
Which is, you know, one percentage point
better than, than the Netflix, model.

258
00:15:38,900 --> 00:15:42,450
Of course,
the million dollar model is down here.

259
00:15:42,450 --> 00:15:44,710
We need to be at 0.85.

260
00:15:44,710 --> 00:15:46,800
And what we will do now
in the rest of the,

261
00:15:46,800 --> 00:15:51,670
of, of the videos on recommender systems,
we will learn how to basically go

262
00:15:51,670 --> 00:15:56,403
all the way down to here, and
how kind of win $1 million.

263
00:15:56,403 --> 00:15:59,290
So right now collaborative
filtering is at 0.94.

264
00:15:59,290 --> 00:16:04,030
The question is, what kind of methods
can we, can we use all the way down and

265
00:16:04,030 --> 00:16:06,720
reduce our error to 0.85?

266
00:16:06,720 --> 00:16:07,682
So that's the idea.

