1
00:00:08,460 --> 00:00:14,360
In this video we are going to talk about support vector machines.

2
00:00:14,360 --> 00:00:18,955
Let's think again about the classification task we first saw.

3
00:00:18,955 --> 00:00:23,415
This is a passage about some medical condition,

4
00:00:23,415 --> 00:00:29,020
and we need to decide whether it is nephrology, neurology, or podiatry.

5
00:00:29,020 --> 00:00:32,840
Looking at the words, like athlete's foot,

6
00:00:32,840 --> 00:00:35,120
infection of the foot and so on,

7
00:00:35,120 --> 00:00:38,965
we would know that this is podiatry.

8
00:00:38,965 --> 00:00:40,380
Let's take another example.

9
00:00:40,380 --> 00:00:43,020
And this is about sentiment analysis.

10
00:00:43,020 --> 00:00:47,025
This is somebody's review for a movie,

11
00:00:47,025 --> 00:00:52,470
and you could guess which movie it is since in sensibility- and you can

12
00:00:52,470 --> 00:00:58,230
see that the the author or the reviewer here has given some stars and so on,

13
00:00:58,230 --> 00:01:02,895
but then used words like wow and great movie and so on.

14
00:01:02,895 --> 00:01:09,415
Using those words we could then identify that this particular review is positive.

15
00:01:09,415 --> 00:01:12,600
If you have seen words like boring, or lame,

16
00:01:12,600 --> 00:01:14,040
or worst, and so on,

17
00:01:14,040 --> 00:01:16,370
it would be a negative review.

18
00:01:16,370 --> 00:01:19,095
So again this is a sentiment classification task,

19
00:01:19,095 --> 00:01:20,550
where looking at the words,

20
00:01:20,550 --> 00:01:24,605
we had to determine whether it's positive or negative.

21
00:01:24,605 --> 00:01:30,880
So in general, a classifier can be said to be a function on the input data.

22
00:01:30,880 --> 00:01:32,775
So you have some function F,

23
00:01:32,775 --> 00:01:37,260
that works on medical passages and labels,

24
00:01:37,260 --> 00:01:40,605
or decides that what the class label should be.

25
00:01:40,605 --> 00:01:43,405
Nephrology, neurology or podiatry.

26
00:01:43,405 --> 00:01:49,245
Or another function looks at this review, a movie review,

27
00:01:49,245 --> 00:01:53,840
and decides it's +1 that stands for positive review,

28
00:01:53,840 --> 00:01:57,135
or -1 that stands for negative review.

29
00:01:57,135 --> 00:01:59,330
It's typical to actually have numbers,

30
00:01:59,330 --> 00:02:02,540
like +1 or -1 Assigned to classes,

31
00:02:02,540 --> 00:02:07,990
and that's what support vector machines or any of that linear classifier, typically use.

32
00:02:07,990 --> 00:02:12,120
When you're talking about support vector machines,

33
00:02:12,120 --> 00:02:15,600
we need to talk about decision boundaries.

34
00:02:15,600 --> 00:02:17,235
What is a decision boundary?

35
00:02:17,235 --> 00:02:19,580
Let's take this example.

36
00:02:19,580 --> 00:02:22,725
There are just two dimensions, X1 and X2,

37
00:02:22,725 --> 00:02:25,475
and there are points,

38
00:02:25,475 --> 00:02:28,650
leader points in this two dimensional in space.

39
00:02:28,650 --> 00:02:32,560
Some of them are marked star, others are marked circle.

40
00:02:32,560 --> 00:02:35,680
Let's call the circle ones as Class A,

41
00:02:35,680 --> 00:02:38,805
and the stars as Class B.

42
00:02:38,805 --> 00:02:44,120
Now you need to somehow make a boundary around one

43
00:02:44,120 --> 00:02:49,290
of the classes to separate that from the other class.

44
00:02:49,290 --> 00:02:51,595
You can use any decision boundary,

45
00:02:51,595 --> 00:02:56,165
you could use one that is a circle,

46
00:02:56,165 --> 00:03:00,080
or you could use one that's a square,

47
00:03:00,080 --> 00:03:02,770
or you could have any other random shape,

48
00:03:02,770 --> 00:03:05,210
in this case let's take a shape of a heart.

49
00:03:05,210 --> 00:03:11,505
A classic fiction function is represented by this decision surface, or decision boundary.

50
00:03:11,505 --> 00:03:18,785
And this choice of what shape you should use is completely dependent on this application.

51
00:03:18,785 --> 00:03:21,085
If you think that points in

52
00:03:21,085 --> 00:03:26,335
the two dimensional base are well represented using a circle, you would use that.

53
00:03:26,335 --> 00:03:29,900
Whenever you use a decision boundary,

54
00:03:29,900 --> 00:03:34,215
the main purpose is when you get a new point that is unlabeled,

55
00:03:34,215 --> 00:03:37,065
depending on the decision boundary you have learned,

56
00:03:37,065 --> 00:03:40,050
you can label it as one of the classes.

57
00:03:40,050 --> 00:03:41,715
So in this particular example,

58
00:03:41,715 --> 00:03:43,660
this point that is unlabeled,

59
00:03:43,660 --> 00:03:45,485
will be marked as a star,

60
00:03:45,485 --> 00:03:55,195
or Class B because that is outside this boundary that was learned around class A.

61
00:03:55,195 --> 00:03:57,100
But when you're using a decision boundary,

62
00:03:57,100 --> 00:04:00,655
there are other factors that should be considered.

63
00:04:00,655 --> 00:04:05,590
And let's take another example where here you have positive and negative points,

64
00:04:05,590 --> 00:04:07,270
so positive is one class,

65
00:04:07,270 --> 00:04:08,665
negative is the other class.

66
00:04:08,665 --> 00:04:13,810
And all the red positives and negatives define this training data that we have.

67
00:04:13,810 --> 00:04:17,085
Looking at this data,

68
00:04:17,085 --> 00:04:20,530
you could learn a boundary like this.

69
00:04:20,530 --> 00:04:23,040
It's an irregular shaped Pentagon,

70
00:04:23,040 --> 00:04:29,285
or in general any other polygon or closed surface that puts all the positive class,

71
00:04:29,285 --> 00:04:32,345
on the positive points, inside this polygon,

72
00:04:32,345 --> 00:04:36,325
and all the negative points are outside.

73
00:04:36,325 --> 00:04:38,910
This is a perfectly great division boundary.

74
00:04:38,910 --> 00:04:45,225
The accuracy on the training data is 100% Every point is labeled correctly.

75
00:04:45,225 --> 00:04:48,040
But then, when we look at the test data,

76
00:04:48,040 --> 00:04:50,840
represented by black positives and negatives here,

77
00:04:50,840 --> 00:04:54,210
you would start seeing that yes in general there are

78
00:04:54,210 --> 00:04:57,930
positive points near positive points that you saw earlier.

79
00:04:57,930 --> 00:05:01,170
In general negative points are near negative points.

80
00:05:01,170 --> 00:05:05,520
But then, because of the way we defined the division boundary,

81
00:05:05,520 --> 00:05:09,115
there are many many more mistakes that we have done here.

82
00:05:09,115 --> 00:05:13,995
There are many many positive points outside of the polygon.

83
00:05:13,995 --> 00:05:17,310
The outside spear is supposed to be negative,

84
00:05:17,310 --> 00:05:22,380
and there are points within this polygon, that are negative.

85
00:05:22,380 --> 00:05:26,840
So in general, in unseen test data,

86
00:05:26,840 --> 00:05:31,090
this decision boundary has not worked very well.

87
00:05:31,090 --> 00:05:33,980
This problem is called data overfitting.

88
00:05:33,980 --> 00:05:38,735
When you have learned a decision boundary that works very well on training data,

89
00:05:38,735 --> 00:05:41,850
but does not generalize to test data.

90
00:05:41,850 --> 00:05:45,345
So instead of a irregular shape boundary,

91
00:05:45,345 --> 00:05:48,480
let's consider a linear boundary or a straight line.

92
00:05:48,480 --> 00:05:49,993
So in a two dimension space,

93
00:05:49,993 --> 00:05:51,285
it is a straight line.

94
00:05:51,285 --> 00:05:52,740
In a three dimensional space.

95
00:05:52,740 --> 00:05:54,075
It would be a plane,

96
00:05:54,075 --> 00:05:59,485
in general and dimensional data representation, it would be a hyperplane.

97
00:05:59,485 --> 00:06:01,995
So in this particular case,

98
00:06:01,995 --> 00:06:05,700
you have a line that separates the positive region,

99
00:06:05,700 --> 00:06:10,610
that's one half of this two dimensional space as positive,

100
00:06:10,610 --> 00:06:12,900
and the other half as negative.

101
00:06:12,900 --> 00:06:14,340
As you can see,

102
00:06:14,340 --> 00:06:17,125
it's not accurate, it's not 100% accurate.

103
00:06:17,125 --> 00:06:20,665
You have three negative points on the positive side,

104
00:06:20,665 --> 00:06:23,960
and you have one positive point on the negative side.

105
00:06:23,960 --> 00:06:26,715
So yes, it has made some mistakes in the training data,

106
00:06:26,715 --> 00:06:29,025
but it's very easy to find,

107
00:06:29,025 --> 00:06:35,510
because you have to basically learn a line that maximally separates these two points,

108
00:06:35,510 --> 00:06:41,025
that that makes as fewer errors as positive, as possible,

109
00:06:41,025 --> 00:06:44,400
and you can see that there is no one line that could have

110
00:06:44,400 --> 00:06:48,515
separated the positives from negatives completely.

111
00:06:48,515 --> 00:06:51,255
So this is as good as you can get.

112
00:06:51,255 --> 00:06:53,695
It's also very easy to evaluate,

113
00:06:53,695 --> 00:06:58,540
because now you have a very simple rule that any point in

114
00:06:58,540 --> 00:07:01,350
this entire half-space and

115
00:07:01,350 --> 00:07:05,600
half-region that's on the positive side is going to be called positive.

116
00:07:05,600 --> 00:07:11,540
And any point on the other side is going to be called negative.

117
00:07:11,540 --> 00:07:19,145
In general, this idea of using a simple classifier or a simple model to fit the data,

118
00:07:19,145 --> 00:07:20,530
Is called Occam's Razor.

119
00:07:20,530 --> 00:07:25,200
And the rule of thumb here is simple models generalize well.

120
00:07:25,200 --> 00:07:29,020
You want to use something that is simple to explain,

121
00:07:29,020 --> 00:07:31,380
that would work very well in the test data.

122
00:07:31,380 --> 00:07:35,140
And as you can see, if you apply the test points on top of this,

123
00:07:35,140 --> 00:07:40,215
you would notice that all test points have been correctly classified.

124
00:07:40,215 --> 00:07:43,665
Except for one positive that's on the negative side.

125
00:07:43,665 --> 00:07:48,440
So the error is really just one point on the test,

126
00:07:48,440 --> 00:07:52,800
and of course there are four points in the training data that were misclassified as well.

127
00:07:52,800 --> 00:07:55,895
But overall this is still a much better model than

128
00:07:55,895 --> 00:07:59,060
this irregular shaped boundary where we had

129
00:07:59,060 --> 00:08:02,830
seen earlier that quite a few positive points have been misclassified,

130
00:08:02,830 --> 00:08:08,860
and two of the negative points were misclassified in the test data.

131
00:08:08,860 --> 00:08:11,670
So how do you find this linear boundary?

132
00:08:11,670 --> 00:08:16,475
You could use a lot of algorithms and methods to do so.

133
00:08:16,475 --> 00:08:19,285
Typically, when you're finding a linear boundary,

134
00:08:19,285 --> 00:08:23,445
you're finding W, or the slope of the line.

135
00:08:23,445 --> 00:08:25,555
You basically want to find out

136
00:08:25,555 --> 00:08:31,966
the coefficients that are the weights associated with the dimensions X1 and X2,

137
00:08:31,966 --> 00:08:34,570
so W is typically of the same dimension,

138
00:08:34,570 --> 00:08:38,415
same size as your data points.

139
00:08:38,415 --> 00:08:41,770
And you are going to use a linear function

140
00:08:41,770 --> 00:08:46,595
to see which point would be called positive or negative.

141
00:08:46,595 --> 00:08:49,525
There could be many methods, perceptron,

142
00:08:49,525 --> 00:08:52,050
a linear discriminant analysis,

143
00:08:52,050 --> 00:08:56,985
or linear least square techniques and so on.

144
00:08:56,985 --> 00:08:58,625
There are many methods to use,

145
00:08:58,625 --> 00:09:00,740
so you could use some perceptrons,

146
00:09:00,740 --> 00:09:03,135
you could use linear discriminative analysis,

147
00:09:03,135 --> 00:09:05,215
or linear least squares.

148
00:09:05,215 --> 00:09:08,280
And in one particular case, in this case,

149
00:09:08,280 --> 00:09:12,035
we're going to talk about support vector machines.

150
00:09:12,035 --> 00:09:15,240
The problem is that whenever you have

151
00:09:15,240 --> 00:09:18,955
a linear boundary that separates the positive points and negative points,

152
00:09:18,955 --> 00:09:24,215
if there is one line that splits the positives and negatives from each other,

153
00:09:24,215 --> 00:09:26,745
there are in fact infinitely many lines.

154
00:09:26,745 --> 00:09:28,665
So you could have a line like this,

155
00:09:28,665 --> 00:09:31,045
or the line like this, or the line like this.

156
00:09:31,045 --> 00:09:34,695
They all separate the positive points from the negative points.

157
00:09:34,695 --> 00:09:38,570
Then the choice becomes what is the best line here.

158
00:09:38,570 --> 00:09:40,700
So let's take another example,

159
00:09:40,700 --> 00:09:42,840
very similar to what we have seen before.

160
00:09:42,840 --> 00:09:46,325
Again two dimensions, positives and negatives.

161
00:09:46,325 --> 00:09:48,600
And consider this line.

162
00:09:48,600 --> 00:09:54,705
This line is one that goes from the farthest positive to the farthest negative.

163
00:09:54,705 --> 00:09:58,815
By farthest I mean the ones that are kind of the most confusing,

164
00:09:58,815 --> 00:10:01,320
that are closest to the boundary,

165
00:10:01,320 --> 00:10:05,510
and this line perfectly splits the data as best as it can.

166
00:10:05,510 --> 00:10:07,065
I mean it's not really 100% accurate,

167
00:10:07,065 --> 00:10:09,300
as you can see that there are negative points on

168
00:10:09,300 --> 00:10:13,045
the positive side and two points on the negative side.

169
00:10:13,045 --> 00:10:16,105
But if you learn this line,

170
00:10:16,105 --> 00:10:19,000
you could actually- you learned any other line.

171
00:10:19,000 --> 00:10:22,535
And the question becomes what is a reasonable boundary?

172
00:10:22,535 --> 00:10:24,820
So instead of the blue line,

173
00:10:24,820 --> 00:10:28,240
if you consider this red band,

174
00:10:28,240 --> 00:10:33,815
you would see that the points that are in the confusion are low.

175
00:10:33,815 --> 00:10:35,775
What do I mean by that?

176
00:10:35,775 --> 00:10:37,300
So in the previous example,

177
00:10:37,300 --> 00:10:39,895
when you had just this blue line,

178
00:10:39,895 --> 00:10:46,315
any small change to the positive point or a negative point will make an error.

179
00:10:46,315 --> 00:10:50,840
If you make a small perturbation on the posture point,

180
00:10:50,840 --> 00:10:52,930
let's take an example here,

181
00:10:52,930 --> 00:10:54,885
and say this point,

182
00:10:54,885 --> 00:11:00,560
is moved a little bit here, suddenly,

183
00:11:00,560 --> 00:11:08,605
you're either making a mistake or this line has to change to now get here.

184
00:11:08,605 --> 00:11:11,560
So you can see that a small change in

185
00:11:11,560 --> 00:11:16,390
a data point can actually lead to a big change in the classifier.

186
00:11:16,390 --> 00:11:21,890
However, if you use a line like this,

187
00:11:21,890 --> 00:11:26,530
where you have a big band separating the points,

188
00:11:26,530 --> 00:11:30,635
you would notice that any small change,

189
00:11:30,635 --> 00:11:35,385
the same small change that you found here,

190
00:11:35,385 --> 00:11:41,855
would still be on the same side as the hyperplane,

191
00:11:41,855 --> 00:11:44,060
as what you would expect it to be.

192
00:11:44,060 --> 00:11:47,390
So a small change in your data points would still not

193
00:11:47,390 --> 00:11:51,695
change the label that this particular classifier will assign that.

194
00:11:51,695 --> 00:12:01,640
So a small change- it's actually more- So it's more resistant to noise or perturbations.

195
00:12:01,640 --> 00:12:04,535
This idea is called maximum margin.

196
00:12:04,535 --> 00:12:09,920
You would want to learn the thickest band that you can fit,

197
00:12:09,920 --> 00:12:13,335
that separates the positive points from negative points.

198
00:12:13,335 --> 00:12:16,250
This band, or the thickness is called the margin,

199
00:12:16,250 --> 00:12:21,165
and the middle line between them is called the actual hyperplane.

200
00:12:21,165 --> 00:12:25,310
So that's the maximum margin hyperplane.

201
00:12:25,310 --> 00:12:31,250
The idea here is that you are not basing your classifier on two points,

202
00:12:31,250 --> 00:12:32,430
a positive and negative point,

203
00:12:32,430 --> 00:12:34,985
but a set of points called support vectors.

204
00:12:34,985 --> 00:12:40,340
And in general, the support vector machine is just maximum margin classifier.

205
00:12:40,340 --> 00:12:44,655
Where it's not just one or two points that will make a big difference,

206
00:12:44,655 --> 00:12:50,795
you actually have a lot of support to learn this margin or this band,

207
00:12:50,795 --> 00:12:56,730
and these support vectors are the points that are the most sensitive to shift.

208
00:12:56,730 --> 00:12:59,840
But even then, small changes or perturbations to

209
00:12:59,840 --> 00:13:03,785
support vectors still would not change the classification,

210
00:13:03,785 --> 00:13:09,845
because classification happens with respect to this maximum margin hyperplane.

211
00:13:09,845 --> 00:13:12,790
So now the question becomes how do you find it.

212
00:13:12,790 --> 00:13:17,240
Support vector machines uses optimization techniques to do it.

213
00:13:17,240 --> 00:13:22,185
These are linear classifiers that find hyperplane to separate these two classes of data,

214
00:13:22,185 --> 00:13:24,005
the positive and the negative class,

215
00:13:24,005 --> 00:13:27,726
and basically given training data of the type X_1,Y_1,

216
00:13:27,726 --> 00:13:29,870
X_2,Y_2 and so on,

217
00:13:29,870 --> 00:13:35,870
where Xs are the representations of the data in terms of its features,

218
00:13:35,870 --> 00:13:38,845
as opposed to any dimensional feature representation.

219
00:13:38,845 --> 00:13:43,115
And you have the Ys as the class label,

220
00:13:43,115 --> 00:13:45,415
which is one of two labels here, -1 and +1,

221
00:13:45,415 --> 00:13:49,385
then SVM will find the linear function W,

222
00:13:49,385 --> 00:13:54,070
a weight vector, such that the function of f, of X_i,

223
00:13:54,070 --> 00:13:58,276
is this dot product of W and X_i,

224
00:13:58,276 --> 00:14:01,895
so W.X_i with some bias done,

225
00:14:01,895 --> 00:14:07,750
but the idea is that if the value for a new X_i.

226
00:14:07,750 --> 00:14:10,370
If the function leads to a positive value,

227
00:14:10,370 --> 00:14:12,455
of X_i is greater than zero,

228
00:14:12,455 --> 00:14:19,290
then the label is +1, otherwise the label is -1.

229
00:14:19,290 --> 00:14:24,475
You have seen that this SVM tend to work only for binary classification problems,

230
00:14:24,475 --> 00:14:25,900
and we have seen an example of this,

231
00:14:25,900 --> 00:14:28,250
+1 and -1 so far,

232
00:14:28,250 --> 00:14:31,320
but then what happens when you have multiple classes?

233
00:14:31,320 --> 00:14:38,845
Let's take this example of the three classes, podiatry, nephrology, neurology.

234
00:14:38,845 --> 00:14:43,545
In general, when you were to use the SVM in a multiclass classification set up,

235
00:14:43,545 --> 00:14:49,620
you would want to learn it in one of two scenarios.

236
00:14:49,620 --> 00:14:53,085
One is called One versus rest.

237
00:14:53,085 --> 00:14:55,410
Where you are actually learning

238
00:14:55,410 --> 00:15:00,950
binary classifiers between one class and all the other classes together.

239
00:15:00,950 --> 00:15:04,420
For example, this is the classifier

240
00:15:04,420 --> 00:15:08,895
between nephrology and all documents that are not nephrology.

241
00:15:08,895 --> 00:15:12,910
Assume that d1 to d9 are all nephrology documents.

242
00:15:12,910 --> 00:15:17,635
So this classification is a perfect classification to learn nephrology,

243
00:15:17,635 --> 00:15:19,395
as these nine documents,

244
00:15:19,395 --> 00:15:23,565
and anything that is not nephrology is on the other side.

245
00:15:23,565 --> 00:15:27,230
Then, we learn another classifier,

246
00:15:27,230 --> 00:15:28,870
this time for neurology.

247
00:15:28,870 --> 00:15:33,335
Where the documents d10 to d15 are neurology documents,

248
00:15:33,335 --> 00:15:35,870
and everything else, d1 to d9,

249
00:15:35,870 --> 00:15:37,865
and d16 to d21,

250
00:15:37,865 --> 00:15:41,030
are all not neurology.

251
00:15:41,030 --> 00:15:44,380
And then, we do the same thing with podiatry.

252
00:15:44,380 --> 00:15:46,970
You have the documents d16 to the d21,

253
00:15:46,970 --> 00:15:49,445
that is labeled podiatry, and everything else.

254
00:15:49,445 --> 00:15:50,970
That's d1 to d15,

255
00:15:50,970 --> 00:15:53,785
is going to be not podiatry.

256
00:15:53,785 --> 00:15:56,650
So in general, we are learning three classifiers,

257
00:15:56,650 --> 00:15:59,100
or for n class setup,

258
00:15:59,100 --> 00:16:00,405
we are learning n classifiers,

259
00:16:00,405 --> 00:16:05,405
such that all the points in one region,

260
00:16:05,405 --> 00:16:10,265
let's take the region as this,

261
00:16:10,265 --> 00:16:17,972
D16 to d21, ist' going to be not nephrology,

262
00:16:17,972 --> 00:16:24,665
yes on podiatry, And what's the other one.

263
00:16:24,665 --> 00:16:27,815
Yes, it's not on neurology.

264
00:16:27,815 --> 00:16:29,130
So it's not neurology,

265
00:16:29,130 --> 00:16:30,905
it's not nephrology, and it's podiatry.

266
00:16:30,905 --> 00:16:33,220
So this is going to be podiatry.

267
00:16:33,220 --> 00:16:38,015
Right? Let's take another set up,

268
00:16:38,015 --> 00:16:41,540
and this set up is where instead of learning one versus rest,

269
00:16:41,540 --> 00:16:45,640
we are going to do one versus one.

270
00:16:45,640 --> 00:16:48,435
That means that you're going to learn classifiers

271
00:16:48,435 --> 00:16:53,090
between let's say nephrology and neurology.

272
00:16:53,090 --> 00:16:57,380
So in this case you look at only documents that were labeled nephrology or neurology.

273
00:16:57,380 --> 00:17:01,870
That's d1 to d9, and d10 to d15, and separate them.

274
00:17:01,870 --> 00:17:04,955
So the red line just separates these two points.

275
00:17:04,955 --> 00:17:07,570
These two classes, I'm sorry.

276
00:17:07,570 --> 00:17:09,590
And then you learn another classifier.

277
00:17:09,590 --> 00:17:12,155
This is between neurology and podiatry,

278
00:17:12,155 --> 00:17:14,810
that works on a smaller set.

279
00:17:14,810 --> 00:17:19,250
Again it's between the d1 to d9 as the neurology documents,

280
00:17:19,250 --> 00:17:23,260
and d16 to d21 as the podiatry document.

281
00:17:23,260 --> 00:17:24,950
And again the third one,

282
00:17:24,950 --> 00:17:29,000
which is between nephrology and podiatry.

283
00:17:29,000 --> 00:17:31,325
When you have these three classes together,

284
00:17:31,325 --> 00:17:33,178
you can see that in general,

285
00:17:33,178 --> 00:17:36,980
for an n-class set up, there are (n,2).

286
00:17:36,980 --> 00:17:43,040
A combonotorial number n_square number of classifiers.

287
00:17:43,040 --> 00:17:45,770
It so happens that for three classes,

288
00:17:45,770 --> 00:17:48,005
it is three classifiers,

289
00:17:48,005 --> 00:17:52,710
But for four classes it'll likely to be six classifiers, and so on.

290
00:17:52,710 --> 00:18:01,900
But then, you can still use the same idea where you have these points,

291
00:18:01,900 --> 00:18:11,785
going to be called between the class labels podiatry and neurology, it's podiatry.

292
00:18:11,785 --> 00:18:15,717
Between the classes podiatry and nephrology.

293
00:18:15,717 --> 00:18:20,770
It's podiatry, and between the classes nephrology and neurology,

294
00:18:20,770 --> 00:18:24,190
it's basically nothing, because it's right in the middle, right?

295
00:18:24,190 --> 00:18:27,565
But you still see that sometimes it's going to be called nephrology,

296
00:18:27,565 --> 00:18:29,260
sometimes it's going to be called neurology,

297
00:18:29,260 --> 00:18:33,970
but the overall votes are still in favor of podiatry because you have

298
00:18:33,970 --> 00:18:36,550
two- that all of these points you're going to get

299
00:18:36,550 --> 00:18:39,925
two votes on podiatry because of the other two classes.

300
00:18:39,925 --> 00:18:44,650
And maybe one additional vote for nephrology or neurology.

301
00:18:44,650 --> 00:18:49,995
In any case, all of these points are still going to be classified correctly as podiatry.

302
00:18:49,995 --> 00:18:51,305
So both of these,

303
00:18:51,305 --> 00:18:55,135
in this particular toy dataset have been labeled the right way,

304
00:18:55,135 --> 00:18:58,390
but just the idea and the approach is very different.

305
00:18:58,390 --> 00:19:02,530
In one case you're learning one class against all the other classes,

306
00:19:02,530 --> 00:19:09,040
and in the other case it's one versus one and do n_square a number of classes.

307
00:19:09,040 --> 00:19:11,850
Or reference squared number of classes this way.

308
00:19:11,850 --> 00:19:17,035
Okay, so let's look at one of the parameters when we learn a support vector machine.

309
00:19:17,035 --> 00:19:20,405
One of the most critical parameters becomes parameter C,

310
00:19:20,405 --> 00:19:23,265
and that defines regularization.

311
00:19:23,265 --> 00:19:26,530
Regularization is a term that denotes how

312
00:19:26,530 --> 00:19:31,790
important it is for individual data point to be labeled correctly,

313
00:19:31,790 --> 00:19:35,440
as compared to all the points in the general model.

314
00:19:35,440 --> 00:19:39,285
So how much importance should we give to individual data points?

315
00:19:39,285 --> 00:19:41,360
This parameter for example,

316
00:19:41,360 --> 00:19:43,625
if you have a larger value of c,

317
00:19:43,625 --> 00:19:46,400
that means the regularization is less,

318
00:19:46,400 --> 00:19:50,955
and that means that you are fitting the training data as much as possible.

319
00:19:50,955 --> 00:19:54,780
You are giving the individual points a lot of importance,

320
00:19:54,780 --> 00:19:57,305
and every data point becomes important.

321
00:19:57,305 --> 00:19:58,925
You want to get them right,

322
00:19:58,925 --> 00:20:01,370
even if the overall accuracy is low,

323
00:20:01,370 --> 00:20:04,070
or generalization error is high.

324
00:20:04,070 --> 00:20:08,042
Whereas smaller values of c means that are more regularization,

325
00:20:08,042 --> 00:20:11,990
where you are tolerant to errors on individual data points,

326
00:20:11,990 --> 00:20:16,435
as long as in the overall scheme you are getting simpler models let's say.

327
00:20:16,435 --> 00:20:21,140
So the generalization error is expected to be low.

328
00:20:21,140 --> 00:20:26,010
The default values in most of the packages is one.

329
00:20:26,010 --> 00:20:29,300
But you could change that parameter depending on how important you

330
00:20:29,300 --> 00:20:34,410
want your individual data point classifications to be.

331
00:20:34,410 --> 00:20:38,285
The other parameters are typically with respect to

332
00:20:38,285 --> 00:20:43,420
what is the type of a decision boundary you would want to learn.

333
00:20:43,420 --> 00:20:46,005
We've talked about linear kernels a lot here.

334
00:20:46,005 --> 00:20:49,517
Linear decision boundaries, but you could have a polynomial kernel,

335
00:20:49,517 --> 00:20:51,940
or a regular basis function kernel and so on.

336
00:20:51,940 --> 00:20:55,535
In fact we have talked about these in course three,

337
00:20:55,535 --> 00:20:57,490
as part of this specialization.

338
00:20:57,490 --> 00:21:00,460
So I would refer you to some of the slides and discussion

339
00:21:00,460 --> 00:21:05,050
around that in the previous course, in the course three.

340
00:21:05,050 --> 00:21:10,735
The other parameter for SVM is whether it's multiclass or not,

341
00:21:10,735 --> 00:21:11,995
and if it is multiclass,

342
00:21:11,995 --> 00:21:13,105
which one would you use,

343
00:21:13,105 --> 00:21:16,720
so OVR would be the parameter setting for one versus rest,

344
00:21:16,720 --> 00:21:19,030
and there are other parameter options possible.

345
00:21:19,030 --> 00:21:20,655
As you can see with one versus rest,

346
00:21:20,655 --> 00:21:23,205
you are learning fewer number of classifiers.

347
00:21:23,205 --> 00:21:26,585
So that's preferred over one versus one.

348
00:21:26,585 --> 00:21:32,420
The other important factor and a parameter would be class weight.

349
00:21:32,420 --> 00:21:35,290
So different classes could get different weights.

350
00:21:35,290 --> 00:21:37,870
For example, if you want a particular class,

351
00:21:37,870 --> 00:21:39,970
let's say spam or not spam,

352
00:21:39,970 --> 00:21:46,745
and you know that the spams are usually like 80% of e-mails somebody gets.

353
00:21:46,745 --> 00:21:47,950
So then he would want,

354
00:21:47,950 --> 00:21:52,960
that because it's just such a skewed distribution where one of the classes 80% and

355
00:21:52,960 --> 00:22:00,025
the other classes 20% you would want to give different weight to these two classes.

356
00:22:00,025 --> 00:22:06,580
In general these these parameters are possible to set in in python models,

357
00:22:06,580 --> 00:22:10,395
and we'll talk about exact parameter settings in another video.

358
00:22:10,395 --> 00:22:15,115
But to conclude, the big take home messages from support vector the machines is that,

359
00:22:15,115 --> 00:22:18,220
these tend to be most accurate classifieds for text,

360
00:22:18,220 --> 00:22:20,740
especially when we are talking about high-dimensional data,

361
00:22:20,740 --> 00:22:23,750
as text data typically is.

362
00:22:23,750 --> 00:22:29,095
There are strong theoretical foundations for support vector machines.

363
00:22:29,095 --> 00:22:31,450
It's based out of optimization theory,

364
00:22:31,450 --> 00:22:35,305
which were not talked about in any detail in this video,

365
00:22:35,305 --> 00:22:39,200
but I would encourage interested folks to go and check it out.

366
00:22:39,200 --> 00:22:43,580
We'll add some of them in the reading list for this video.

367
00:22:43,580 --> 00:22:46,275
Support vector machine handles only numeric data.

368
00:22:46,275 --> 00:22:47,905
So what would you do with other data?

369
00:22:47,905 --> 00:22:51,500
You typically can word these categorical features into numeric features.

370
00:22:51,500 --> 00:22:55,820
As you can see this is a dot product based algorithm.

371
00:22:55,820 --> 00:23:01,445
So it uses numbers to define whether it should be boundary on one side or the other.

372
00:23:01,445 --> 00:23:05,230
So it works only on numeric features.

373
00:23:05,230 --> 00:23:08,140
You would also want to normalize the features.

374
00:23:08,140 --> 00:23:12,100
That means you don't want some dimensions to be very high in numbers,

375
00:23:12,100 --> 00:23:14,505
and other dimensions to be very low.

376
00:23:14,505 --> 00:23:17,710
You would want to put them all in a zero to one range,

377
00:23:17,710 --> 00:23:19,615
so that each feature gets

378
00:23:19,615 --> 00:23:23,910
enough importance and you can learn the relative importance that way.

379
00:23:23,910 --> 00:23:27,710
But in general, the hyperplane that you learn even that

380
00:23:27,710 --> 00:23:31,130
they're very easy to learn and easy to test,

381
00:23:31,130 --> 00:23:32,780
they are hard to interpret.

382
00:23:32,780 --> 00:23:34,730
Which means that you don't really know

383
00:23:34,730 --> 00:23:38,805
why a particular point has been labelled as positive or negative.

384
00:23:38,805 --> 00:23:41,540
It's a combination of multiple factors.

385
00:23:41,540 --> 00:23:48,199
But in general if explanation is not a critical requirement for the classifier,

386
00:23:48,199 --> 00:23:52,195
support vector machines do give you very high and very accurate classifiers,

387
00:23:52,195 --> 00:23:53,989
very similar to naive bayes.

388
00:23:53,989 --> 00:23:57,400
In fact for a lot of text classification problems,

389
00:23:57,400 --> 00:24:00,120
support vector machines should be one of the first ones you should try.