1
00:00:07,880 --> 00:00:13,700
Okay. So now that we know what is a theoretical understanding of text classification,

2
00:00:13,700 --> 00:00:17,345
let's see how to build one in Python.

3
00:00:17,345 --> 00:00:22,120
There are quite a few toolkits available for supervised text classification.

4
00:00:22,120 --> 00:00:25,215
Scikit-learn is one of them.

5
00:00:25,215 --> 00:00:29,190
For those of you who have gone through course three of the specialization,

6
00:00:29,190 --> 00:00:31,695
you have seen scikit-learn before.

7
00:00:31,695 --> 00:00:35,200
The other toolkit is something that we have seen in this course.

8
00:00:35,200 --> 00:00:38,160
That's NLTK. And in fact,

9
00:00:38,160 --> 00:00:41,385
NLTK interfaces with scikit- learn

10
00:00:41,385 --> 00:00:46,400
and also has interfaces for other machine learning toolkits like Weka.

11
00:00:46,400 --> 00:00:49,615
We will not be covering Weka in this course,

12
00:00:49,615 --> 00:00:53,885
but I would encourage you to check it out as well.

13
00:00:53,885 --> 00:00:56,908
Let's first start with scikit-learn.

14
00:00:56,908 --> 00:01:01,400
Scikit-learn is an open-source machine learning library.

15
00:01:01,400 --> 00:01:05,320
It started as a Google summer of code in 2007,

16
00:01:05,320 --> 00:01:08,275
but it has very strong programmatic interface,

17
00:01:08,275 --> 00:01:10,450
especially, when you compare it against Weka which is

18
00:01:10,450 --> 00:01:16,375
a more graphical user interface or primarily driven as the user interface model.

19
00:01:16,375 --> 00:01:22,020
Scikit-learn is used extensively as a machine learning library in Python.

20
00:01:22,020 --> 00:01:26,605
Scikit-learn has predefined classifiers.

21
00:01:26,605 --> 00:01:29,700
The algorithms are already there for you to use.

22
00:01:29,700 --> 00:01:33,990
So you could use the Naive Bayes Classifier if you want to learn that.

23
00:01:33,990 --> 00:01:37,840
So let's go through some steps about what functions you'd use,

24
00:01:37,840 --> 00:01:39,070
what calls you'd use,

25
00:01:39,070 --> 00:01:42,011
when you're using the Naive Bayes classifier.

26
00:01:42,011 --> 00:01:46,750
First, you need to import Naive Bayes from sklearn.

27
00:01:46,750 --> 00:01:51,610
Then, you're going to call this naive_bayes.MultinomialNB()=clfr

28
00:01:51,610 --> 00:01:56,995
and that would be your Bayes classifier.

29
00:01:56,995 --> 00:01:59,800
Now we have seen earlier that there are

30
00:01:59,800 --> 00:02:03,905
two big ways in which Naive Bayes models can be trained.

31
00:02:03,905 --> 00:02:05,580
One is a multinomial model,

32
00:02:05,580 --> 00:02:07,670
other one is a Bernoulli model.

33
00:02:07,670 --> 00:02:09,880
And you also have a Bernoulli model here,

34
00:02:09,880 --> 00:02:14,590
you have naive_bayes.BernoulliNB() if you want to use that model.

35
00:02:14,590 --> 00:02:18,505
Once you have defined a Bayes classifier,

36
00:02:18,505 --> 00:02:22,010
you can train that classifier on the training data.

37
00:02:22,010 --> 00:02:31,230
You would use classifier.fit and then give the data and label as a two parameters.

38
00:02:31,230 --> 00:02:35,100
If you are completely comfortable with this,

39
00:02:35,100 --> 00:02:37,460
you could actually merge them two together.

40
00:02:37,460 --> 00:02:45,295
You could say naive_bayes.MultinomialNB.fit(train_data, train_labels).

41
00:02:45,295 --> 00:02:47,580
Once you have trained the model,

42
00:02:47,580 --> 00:02:53,805
you would predict the label for a new dataset using predict function.

43
00:02:53,805 --> 00:03:00,639
So you'll have classify that predict and pass the test data that has been get of,

44
00:03:00,639 --> 00:03:04,450
you have already extracted the features and so on.

45
00:03:04,450 --> 00:03:07,170
And then, when you call it with predict,

46
00:03:07,170 --> 00:03:09,510
you're going to be at the labels.

47
00:03:09,510 --> 00:03:11,560
And then once you have the labels,

48
00:03:11,560 --> 00:03:13,960
you can see how well you have done in your classification.

49
00:03:13,960 --> 00:03:16,300
Especially, if you have label test data,

50
00:03:16,300 --> 00:03:18,250
you would use metrics.f1_score,

51
00:03:18,250 --> 00:03:24,150
that is one of the measures that be use and give the test labels.

52
00:03:24,150 --> 00:03:26,760
That is the goal set, the actual labels,

53
00:03:26,760 --> 00:03:29,680
the predicted labels and define

54
00:03:29,680 --> 00:03:33,855
what kind of averaging you want to do: micro averaging and macro averaging.

55
00:03:33,855 --> 00:03:37,980
You have covered some of these concepts already in course three,

56
00:03:37,980 --> 00:03:39,395
so I'll not repeat them here.

57
00:03:39,395 --> 00:03:42,335
But I would point to some of the reading material

58
00:03:42,335 --> 00:03:46,333
around: what kind of measures you could use instead of f1_score,

59
00:03:46,333 --> 00:03:49,725
what does f1 mean and what kind of averaging you could do,

60
00:03:49,725 --> 00:03:54,120
like micro averaging and macro averaging and so on.

61
00:03:54,120 --> 00:03:59,820
Sklearn, that's Scikit-learn, also has SVM classifier.

62
00:03:59,820 --> 00:04:04,900
So how do you train as support vector machine or SVM?

63
00:04:04,900 --> 00:04:07,660
Well, the goals are very similar.

64
00:04:07,660 --> 00:04:14,813
In this case, you are going to import SVM from sklearn and call svm.

65
00:04:14,813 --> 00:04:17,010
SVC as the classifier.

66
00:04:17,010 --> 00:04:20,340
SVC stands for support vector classifier.

67
00:04:20,340 --> 00:04:26,855
As we see, you need to pass some parameters.

68
00:04:26,855 --> 00:04:29,040
Typically, for text classification models,

69
00:04:29,040 --> 00:04:32,080
you're going to focus on linear classifiers.

70
00:04:32,080 --> 00:04:35,960
So it's a linear kernel to consist kernel as linear.

71
00:04:35,960 --> 00:04:39,035
And then you can specify the C parameter.

72
00:04:39,035 --> 00:04:40,840
We have talked about that earlier in one of

73
00:04:40,840 --> 00:04:46,275
the earlier videos and that is the parameter for soft margin.

74
00:04:46,275 --> 00:04:49,110
The default values for kernel is RBF,

75
00:04:49,110 --> 00:04:51,215
a radial basis function,

76
00:04:51,215 --> 00:04:55,545
kernel and the default value for C is one,

77
00:04:55,545 --> 00:05:00,705
where you are neither too hard not too soft on the margin.

78
00:05:00,705 --> 00:05:02,490
Once you have defined this bayes classifier,

79
00:05:02,490 --> 00:05:07,160
you can fit or you can train it the same way as you train the Naive Bayes one.

80
00:05:07,160 --> 00:05:09,765
So you're going to see classifier.fit

81
00:05:09,765 --> 00:05:13,880
and pass the data and labels as two separate parameters.

82
00:05:13,880 --> 00:05:17,210
And then you'll be able to predict the same way as last time you have

83
00:05:17,210 --> 00:05:20,820
classified or predict on test data.

84
00:05:20,820 --> 00:05:23,920
Now, we need to talk briefly about model selection.

85
00:05:23,920 --> 00:05:29,720
You'd recall that there were multiple phases in supervised learning task.

86
00:05:29,720 --> 00:05:31,230
And we talked about it earlier.

87
00:05:31,230 --> 00:05:33,575
There's a training phase and an inference phase.

88
00:05:33,575 --> 00:05:37,870
You have labeled it as that is already labeled,

89
00:05:37,870 --> 00:05:39,815
so in this case, it's green and red.

90
00:05:39,815 --> 00:05:43,175
And you split that label data into

91
00:05:43,175 --> 00:05:46,589
the training data set and the hold out validation data set.

92
00:05:46,589 --> 00:05:50,660
And then you have the test data that could also be

93
00:05:50,660 --> 00:05:56,320
labeled that would be used to say how well you have performed on unseen data.

94
00:05:56,320 --> 00:05:59,580
But typically, the test data set is not labeled.

95
00:05:59,580 --> 00:06:02,030
So you're going to train something on

96
00:06:02,030 --> 00:06:06,665
a label set and then you need to apply it on an unlabeled tests.

97
00:06:06,665 --> 00:06:15,385
So, you need to use some of the labeled data to see how well these models are.

98
00:06:15,385 --> 00:06:17,920
Especially if you're comparing between models,

99
00:06:17,920 --> 00:06:21,700
or if you are tuning a model.

100
00:06:21,700 --> 00:06:22,945
So if you have some parameters,

101
00:06:22,945 --> 00:06:25,480
for example you have the C parameter in SVM,

102
00:06:25,480 --> 00:06:28,840
you need to know what is a good value of C. So,

103
00:06:28,840 --> 00:06:30,220
how would you do it?

104
00:06:30,220 --> 00:06:33,035
That problem is called the model selection problem.

105
00:06:33,035 --> 00:06:34,390
And while you're training,

106
00:06:34,390 --> 00:06:38,590
you need to somehow make sure that you have ways to do that.

107
00:06:38,590 --> 00:06:40,790
There are two ways you could do model selection.

108
00:06:40,790 --> 00:06:49,980
One is keeping some part of printing the label data set separate as the hold out data.

109
00:06:49,980 --> 00:06:55,145
And the other option is cross-validation.

110
00:06:55,145 --> 00:06:56,530
So for the first one,

111
00:06:56,530 --> 00:06:59,140
if you're doing that in scikit-learn,

112
00:06:59,140 --> 00:07:02,620
you're going to save from scikit-learn input model selection.

113
00:07:02,620 --> 00:07:06,285
So that will give you the options available to you.

114
00:07:06,285 --> 00:07:12,105
And then, first we'll see how you could use that train test split.

115
00:07:12,105 --> 00:07:16,675
So I'm going to say model_selection.train_test_split.

116
00:07:16,675 --> 00:07:24,510
Give that train that untrained labels and then specify how much should be your test size.

117
00:07:24,510 --> 00:07:29,050
So for example, suppose you have these data points.

118
00:07:29,050 --> 00:07:33,810
In this case, I have 15 of them and I say I want to do a two third one third split.

119
00:07:33,810 --> 00:07:38,410
So my test size is one third or 0.333.

120
00:07:38,410 --> 00:07:44,235
That would mean 10 of them would be the train set and five of them would be the test.

121
00:07:44,235 --> 00:07:47,905
Now, you could shuffle the training data, the label data,

122
00:07:47,905 --> 00:07:55,665
so that you have a randomly uniform distribution around the positive and negative class.

123
00:07:55,665 --> 00:08:00,110
But then, you could say I wanted to keep let's say

124
00:08:00,110 --> 00:08:04,629
66 percent in train and 33 percent in test or go 80 20 if you want to.

125
00:08:04,629 --> 00:08:10,066
Let's say four out of five should go in my train set and the one out of five,

126
00:08:10,066 --> 00:08:13,460
the fifth part as a test.

127
00:08:13,460 --> 00:08:15,840
When you do it this way,

128
00:08:15,840 --> 00:08:21,930
you are losing out a significant portion of your training data into test.

129
00:08:21,930 --> 00:08:26,845
Remember, that you cannot see the test data when you're training the model.

130
00:08:26,845 --> 00:08:27,985
So test data is used,

131
00:08:27,985 --> 00:08:30,915
exclusively, to tune the parameters.

132
00:08:30,915 --> 00:08:36,975
So your training data effectively reduces to 66 percent in this case.

133
00:08:36,975 --> 00:08:42,245
The other way to do model selection would be cross-validation.

134
00:08:42,245 --> 00:08:45,550
So the cross validation with five full cross-validation,

135
00:08:45,550 --> 00:08:51,566
would be something like this where you split the data into five parts.

136
00:08:51,566 --> 00:08:54,170
These are five folds.

137
00:08:54,170 --> 00:08:59,755
And then, you train five times basically.

138
00:08:59,755 --> 00:09:03,490
You train every time where four parts

139
00:09:03,490 --> 00:09:08,285
are in the train set and one part is in the test set.

140
00:09:08,285 --> 00:09:11,565
So you're going to train five models. Let's see.

141
00:09:11,565 --> 00:09:19,410
First, you're going to train on parts one to four and test on five.

142
00:09:19,410 --> 00:09:24,450
The next time you're going to say I'm going to train on

143
00:09:24,450 --> 00:09:30,270
two to five and test on one and so on.

144
00:09:30,270 --> 00:09:35,090
So you have one iteration where five

145
00:09:35,090 --> 00:09:41,205
is the test and the regression where one is the test,

146
00:09:41,205 --> 00:09:47,065
a third iteration where two is the test and so on.

147
00:09:47,065 --> 00:09:50,610
So when you're doing it in this way,

148
00:09:50,610 --> 00:09:56,390
you get five ways of splitting the data.

149
00:09:56,390 --> 00:10:01,280
Every data point isn't the test ones in this five folds.

150
00:10:01,280 --> 00:10:04,895
And then, you get average out

151
00:10:04,895 --> 00:10:11,475
the five results you get on the whole test set to see how we'll perform,

152
00:10:11,475 --> 00:10:13,430
how the model performs on unseen data.

153
00:10:13,430 --> 00:10:18,730
The cross-validation folds is a parameter.

154
00:10:18,730 --> 00:10:22,255
In this explanation, I took it as five.

155
00:10:22,255 --> 00:10:24,040
It's fairly common to use

156
00:10:24,040 --> 00:10:27,730
10-fold cross-validation especially when you have a large data set.

157
00:10:27,730 --> 00:10:30,670
You can keep 90 percent for training and 10 percent as

158
00:10:30,670 --> 00:10:34,945
the cross validation hold out data set but because you're doing it 10 times,

159
00:10:34,945 --> 00:10:38,830
you're also averaging on multiple runs.

160
00:10:38,830 --> 00:10:42,955
And in fact, it's fairly common to run cross-validation multiple times.

161
00:10:42,955 --> 00:10:47,405
So that you have reduced variance in your results.

162
00:10:47,405 --> 00:10:52,970
Both these models are trained to split and cross-validation are fairly commonly

163
00:10:52,970 --> 00:10:59,515
used and critical when you're doing any model selection. Okay.

164
00:10:59,515 --> 00:11:02,225
Now let's move to NLTK.

165
00:11:02,225 --> 00:11:04,910
How do you do supervised text classification in

166
00:11:04,910 --> 00:11:09,965
the natural language toolkit that we have seen in fair detail in this course?

167
00:11:09,965 --> 00:11:12,960
NLTK has some text classification algorithms.

168
00:11:12,960 --> 00:11:15,980
So for example, it has a naive bayes classifier.

169
00:11:15,980 --> 00:11:17,750
It also has decision trees and

170
00:11:17,750 --> 00:11:22,215
condition exponential models and maximum entropy models and so on.

171
00:11:22,215 --> 00:11:26,090
But the real interesting thing is it has something called

172
00:11:26,090 --> 00:11:31,705
Weka classifier or Sklearn classifier that gives uses of

173
00:11:31,705 --> 00:11:35,360
NLTK a way to call the

174
00:11:35,360 --> 00:11:38,495
underlying scikit-learn classifier or

175
00:11:38,495 --> 00:11:41,480
underlying Weka classifier through their code in Phyton.

176
00:11:41,480 --> 00:11:49,820
Specifically, if you are using the naive bayes classifier that is available in NLTK,

177
00:11:49,820 --> 00:11:55,625
we are going to say from nltk.classify import NaiveBayesClassifier.

178
00:11:55,625 --> 00:11:59,975
You're going to say the classifier is now naivebayesclassifier.train.

179
00:11:59,975 --> 00:12:02,430
So you're directly going to train on the train set to

180
00:12:02,430 --> 00:12:05,580
know that there are no two functions really as

181
00:12:05,580 --> 00:12:08,220
common as I could learn where you have

182
00:12:08,220 --> 00:12:11,840
a based model and then you have a training function.

183
00:12:11,840 --> 00:12:15,090
Here, you are going to say that you have naivebayesclassifier.train and you train

184
00:12:15,090 --> 00:12:20,040
this model and you're going to classify it using the classify function.

185
00:12:20,040 --> 00:12:23,690
So it's classifier.classify (unlabeled instance).

186
00:12:23,690 --> 00:12:27,135
If it's one instance you're going to use the classify function,

187
00:12:27,135 --> 00:12:28,755
if there are many,

188
00:12:28,755 --> 00:12:30,565
I would going to say classify many,

189
00:12:30,565 --> 00:12:34,115
and give a set of unlabeled instances.

190
00:12:34,115 --> 00:12:39,720
You also get the accuracy of the performance of the sklearn classifier using

191
00:12:39,720 --> 00:12:43,680
nltk.classify util function and then

192
00:12:43,680 --> 00:12:48,440
call the accuracy function there where you're giving in the classifier and the test set.

193
00:12:48,440 --> 00:12:50,100
And that will give you how well,

194
00:12:50,100 --> 00:12:54,295
what is the accuracy of this classifier that you have trained.

195
00:12:54,295 --> 00:12:57,450
You can also use other utility functions like labels,

196
00:12:57,450 --> 00:12:59,750
classifier.labels tells you all the labels

197
00:12:59,750 --> 00:13:02,425
that are there that this classifier has trained on.

198
00:13:02,425 --> 00:13:06,920
And you can use some features like this where you have

199
00:13:06,920 --> 00:13:10,400
show_most_informative_features that gives you

200
00:13:10,400 --> 00:13:14,525
the top few features and again said how many features you want.

201
00:13:14,525 --> 00:13:17,885
Say top five or top 10 features that are most

202
00:13:17,885 --> 00:13:23,030
important or informative for the classification task.

203
00:13:23,030 --> 00:13:25,565
It's especially useful in naive bayes classifiers,

204
00:13:25,565 --> 00:13:30,565
when you want to know which features have the most information in them.

205
00:13:30,565 --> 00:13:33,960
Which ones are most informative for the following classifier.

206
00:13:33,960 --> 00:13:35,955
For support vector machines,

207
00:13:35,955 --> 00:13:39,445
there is no native NLTK function.

208
00:13:39,445 --> 00:13:45,435
But as I said, you can use the scikit-learn SVM function through NLTK.

209
00:13:45,435 --> 00:13:48,150
So here, you're going to say nltk.classify

210
00:13:48,150 --> 00:13:51,830
import scikit-learn classifier so SklearnClassifier.

211
00:13:51,830 --> 00:13:58,080
And then, you're going to actually use both naive bayes models from scikit-learn

212
00:13:58,080 --> 00:14:04,860
using a sklearn.naive_bayes import MultinomialNB or BernoulliNB.

213
00:14:04,860 --> 00:14:08,850
And you can use the SVM model that you have seen earlier.

214
00:14:08,850 --> 00:14:13,105
So from sklearn.svm import svc.

215
00:14:13,105 --> 00:14:17,780
You'll call the function very similar way to how you do it in scikit-learn.

216
00:14:17,780 --> 00:14:20,810
And so you have a sklearn.classifier,

217
00:14:20,810 --> 00:14:22,975
give the classifier there,

218
00:14:22,975 --> 00:14:24,515
the name of the classifier,

219
00:14:24,515 --> 00:14:28,800
and.train and then give the train set.

220
00:14:28,800 --> 00:14:33,020
Now for MultinomialNB, there was no parameters that you need to pass.

221
00:14:33,020 --> 00:14:34,790
That's okay. But for support vector machine,

222
00:14:34,790 --> 00:14:36,675
there is one, right?

223
00:14:36,675 --> 00:14:39,350
So you need to specify the kernel for example.

224
00:14:39,350 --> 00:14:43,940
So you can specify that inside this sklearn classifier function.

225
00:14:43,940 --> 00:14:46,940
You're saying that I'm going to call as we see and I'm going to pass

226
00:14:46,940 --> 00:14:50,820
parameters where it's linear kernel and the C parameter,

227
00:14:50,820 --> 00:14:53,255
for example, can also be specified here.

228
00:14:53,255 --> 00:14:57,115
And then you are going to see.train(train_set).

229
00:14:57,115 --> 00:15:01,550
The rest are very similar to how you would do in sklearn.

230
00:15:01,550 --> 00:15:04,910
You're going to use the predict function and so on.

231
00:15:04,910 --> 00:15:07,895
So, the take home messages here are that:

232
00:15:07,895 --> 00:15:11,703
scikit-learn is most commonly use machine learning toolkit in Python,

233
00:15:11,703 --> 00:15:16,260
but NLTK has its own implementation of naive Bayes and it has

234
00:15:16,260 --> 00:15:22,710
this way to interface with scikit-learn and other machine learning toolkits like Weka,

235
00:15:22,710 --> 00:15:25,371
by which you can call those functions,

236
00:15:25,371 --> 00:15:28,000
those implementations through NLTK.