1
00:00:08,100 --> 00:00:16,750
In this video, we are going to talk about basic NLP tasks and introduce you to NLTK.

2
00:00:16,750 --> 00:00:17,951
So what is NLTK?

3
00:00:17,951 --> 00:00:21,170
NLTK stands for Natural Language Toolkit.

4
00:00:21,170 --> 00:00:23,885
It is an open source library in Python,

5
00:00:23,885 --> 00:00:28,100
and we're going to use it extensively in this video and the next.

6
00:00:28,100 --> 00:00:32,060
The advantage of NLTK is that it has support for

7
00:00:32,060 --> 00:00:37,540
most NLP tasks and also provides access to numerous text corpora.

8
00:00:37,540 --> 00:00:43,460
So let's set it up. We first get NLTK in using the import statement,

9
00:00:43,460 --> 00:00:52,225
you have import NLTK and then we can download the text corpora using nltk.download.

10
00:00:52,225 --> 00:00:54,465
It's going to take a little while,

11
00:00:54,465 --> 00:01:01,460
but then once it comes back you can issue a command like this from nltk.book import

12
00:01:01,460 --> 00:01:08,758
* and then it's going to show you the corpora that it has downloaded and made available.

13
00:01:08,758 --> 00:01:11,475
You can see that there are nine text corpora.

14
00:01:11,475 --> 00:01:13,940
Text1 stands for Moby Dick,

15
00:01:13,940 --> 00:01:15,950
text2 is Sense and Sensibility,

16
00:01:15,950 --> 00:01:20,102
you have a Wall Street Journal corpus in text7,

17
00:01:20,102 --> 00:01:24,410
you have some Personals in text8 and Chat Corpus in text5.

18
00:01:24,410 --> 00:01:26,590
So it is quite diverse here.

19
00:01:26,590 --> 00:01:31,240
So as I said text1 is Moby Dick.

20
00:01:31,240 --> 00:01:32,635
If you look at sentences,

21
00:01:32,635 --> 00:01:39,980
it will show you one sentence each from these nine text corpora and sentence one,

22
00:01:39,980 --> 00:01:42,835
Call me Ishmael is from text1.

23
00:01:42,835 --> 00:01:46,040
And then, if you look at how sentence one looks,

24
00:01:46,040 --> 00:01:49,330
sent1 and you see that it has four words.

25
00:01:49,330 --> 00:01:51,620
Call me Ishmael and then full stop.

26
00:01:51,620 --> 00:01:59,025
Now that we have access to text corpus and multiple text corpora,

27
00:01:59,025 --> 00:02:02,385
we can look at counting the vocabulary of words.

28
00:02:02,385 --> 00:02:04,800
So text7 if you recall was

29
00:02:04,800 --> 00:02:11,455
Wall Street Journal and sent7 which was one sentence from text7 is this,

30
00:02:11,455 --> 00:02:13,398
[ 'Pierre', 'Vinken', '61', 'years',

31
00:02:13,398 --> 00:02:15,464
'old', ',', 'will', 'join', 'the', 'board', 'as',

32
00:02:15,464 --> 00:02:20,275
'a', 'nonexecutive', 'director', 'Nov.', '29', '.'

33
00:02:20,275 --> 00:02:22,975
] You have already these words passed out.

34
00:02:22,975 --> 00:02:24,240
So you have comma, separate,

35
00:02:24,240 --> 00:02:25,815
and you have full stop separate.

36
00:02:25,815 --> 00:02:32,810
So the length of sent7 is the number of tokens in this sentence and that's 18.

37
00:02:32,810 --> 00:02:38,320
But if you look at length of text7 that's the entire text corpus,

38
00:02:38,320 --> 00:02:45,300
you'll see that Wall Street Journal has 100,676 words.

39
00:02:45,300 --> 00:02:47,970
It's clear that not all of these are unique.

40
00:02:47,970 --> 00:02:51,058
We can see in the previous example that comma

41
00:02:51,058 --> 00:02:57,045
is repeated twice and full stop is there and words such as "the" and "a" and so on,

42
00:02:57,045 --> 00:03:04,190
are so frequent that they are going to take up a bunch of words from this 100,000 count.

43
00:03:04,190 --> 00:03:09,120
So if you see the unique number of words using the command length of

44
00:03:09,120 --> 00:03:14,480
set of text7 you'll get 12,408.

45
00:03:14,480 --> 00:03:18,180
That means that Wall Street Journal corpus has really only 12,400

46
00:03:18,180 --> 00:03:23,640
unique words even though it is a 100,000-word corpus.

47
00:03:23,640 --> 00:03:26,890
Now that we know how to count words,

48
00:03:26,890 --> 00:03:32,660
let's look at these words and understand how to get the individual frequencies.

49
00:03:32,660 --> 00:03:37,580
So if you want to type out the first 10 words from this set,

50
00:03:37,580 --> 00:03:45,365
first 10 unique words, you'll say, list(set(text7))[:10].

51
00:03:45,365 --> 00:03:47,445
That would give you the first 10 words.

52
00:03:47,445 --> 00:03:48,770
And in this corpus,

53
00:03:48,770 --> 00:03:51,530
the first 10 words really in the set are,

54
00:03:51,530 --> 00:03:54,830
'Mortimer' and 'foul' and 'heights' and 'four' and so on.

55
00:03:54,830 --> 00:04:00,720
You can notice that there is this 'u' and a quote before each word.

56
00:04:00,720 --> 00:04:02,750
Do you recall what it stands for?

57
00:04:02,750 --> 00:04:11,785
You'd recall from the previous videos that 'u' here stands for the UTF-8 encoding.

58
00:04:11,785 --> 00:04:14,815
So these have been automatically UTF-8 encoded.

59
00:04:14,815 --> 00:04:20,260
So each token is represented as a UTF-8 string.

60
00:04:20,260 --> 00:04:23,215
Now, if you have to find out frequency of words,

61
00:04:23,215 --> 00:04:29,216
you're going to use this command, frequency distribution,

62
00:04:29,216 --> 00:04:34,725
FreqDist and then you create this frequency distribution from text7 that is

63
00:04:34,725 --> 00:04:37,375
the Wall Street Journal corpus and store it in

64
00:04:37,375 --> 00:04:43,215
this variable called "Dist" you can start finding statistics from this data structure.

65
00:04:43,215 --> 00:04:47,260
So you have length of Dist and that will give you 12,408.

66
00:04:47,260 --> 00:04:51,145
These are the set of unique words in this word corpus,

67
00:04:51,145 --> 00:04:53,490
this Wall Street Journal Corpus.

68
00:04:53,490 --> 00:04:58,565
Then, you have dist.keys that gives you the actual words.

69
00:04:58,565 --> 00:05:01,010
And that would be your vocab1.

70
00:05:01,010 --> 00:05:03,810
And then if you take the first 10 words of vocab1,

71
00:05:03,810 --> 00:05:08,620
you will get the same 10 words as we saw up there in the top of the slide.

72
00:05:08,620 --> 00:05:14,755
And then, if you want to find out how many times a particular word occurs, you can say,

73
00:05:14,755 --> 00:05:19,126
"Give me the distribution

74
00:05:19,126 --> 00:05:23,600
of this word four," that is UTF encoded and I'll get the response of 20.

75
00:05:23,600 --> 00:05:25,585
That means in this Wall Street Journal corpus,

76
00:05:25,585 --> 00:05:28,870
you have four appearing 20 times.

77
00:05:28,870 --> 00:05:32,425
What if you want to find out how many times

78
00:05:32,425 --> 00:05:37,685
a particular word occurs and also have a condition on the length of the word.

79
00:05:37,685 --> 00:05:42,526
So if you have to find out frequent words and say that I would call a word as

80
00:05:42,526 --> 00:05:48,880
frequent if that word is at least length five and occurs at least a hundred times,

81
00:05:48,880 --> 00:05:53,915
then I can use this command saying w for w in vocab1 if

82
00:05:53,915 --> 00:05:59,745
length of w is greater than five and dist of w is greater than 100.

83
00:05:59,745 --> 00:06:03,970
And then, I'll get this list of words that satisfy

84
00:06:03,970 --> 00:06:06,520
both conditions and you see million and market and

85
00:06:06,520 --> 00:06:10,815
president and trading are the words that satisfy this.

86
00:06:10,815 --> 00:06:14,910
Why did we have a restriction on length of the word?

87
00:06:14,910 --> 00:06:22,010
Because if you don't then words like the or comma or full stop are going to be

88
00:06:22,010 --> 00:06:24,919
very very frequent and those will occur

89
00:06:24,919 --> 00:06:28,880
more than 100 times and they would come up as frequent words.

90
00:06:28,880 --> 00:06:31,115
So this is one way to say, "Oh, you know,

91
00:06:31,115 --> 00:06:35,000
the real unique words are ones that are fairly long,

92
00:06:35,000 --> 00:06:38,360
at least five characters and occurs fairly often."

93
00:06:38,360 --> 00:06:40,331
There are, of course, other ways to do that.

94
00:06:40,331 --> 00:06:43,235
Now, if you look at the next task.

95
00:06:43,235 --> 00:06:44,975
So we know how to count words,

96
00:06:44,975 --> 00:06:47,125
how to find unique words.

97
00:06:47,125 --> 00:06:50,720
The next task becomes normalizing and stemming words.

98
00:06:50,720 --> 00:06:56,420
What does that mean? Normalization is when you have to transform a word to

99
00:06:56,420 --> 00:07:02,720
make it appear the same way or the count even though they look very different.

100
00:07:02,720 --> 00:07:07,840
So for example, there might be different forms in which the same word occurs.

101
00:07:07,840 --> 00:07:14,570
Let's take this example of input1 that has a word list in different forms.

102
00:07:14,570 --> 00:07:16,160
You have it capitalized.

103
00:07:16,160 --> 00:07:18,170
You have it plural, lists,

104
00:07:18,170 --> 00:07:24,960
you have listings and listings and listed as a verb in the past tense and so on.

105
00:07:24,960 --> 00:07:31,070
So the first thing you would want to do is to lowercase them.

106
00:07:31,070 --> 00:07:38,860
Why? Because you don't want to distinguish the capital list with small case list.

107
00:07:38,860 --> 00:07:42,325
So lower would bring it all to lowercase.

108
00:07:42,325 --> 00:07:44,570
And then if you split it on space,

109
00:07:44,570 --> 00:07:47,285
you'll get five words, list,

110
00:07:47,285 --> 00:07:51,110
listed, lists, listing, and listings.

111
00:07:51,110 --> 00:07:52,380
So that was normalization.

112
00:07:52,380 --> 00:07:56,180
Then, comes stemming.

113
00:07:56,180 --> 00:08:05,100
Stemming is to find the root word or the root form of any given word.

114
00:08:05,100 --> 00:08:08,485
You can have multiple algorithms to do stemming.

115
00:08:08,485 --> 00:08:11,971
The ones that are quite popular and used widely is

116
00:08:11,971 --> 00:08:16,240
Porter stemmer and NLTK gives you access to that.

117
00:08:16,240 --> 00:08:21,840
nltk.PorterStemmer would create a stemmer and we call it Porter.

118
00:08:21,840 --> 00:08:26,005
And then, if you stem a word using the Porter stemmer,

119
00:08:26,005 --> 00:08:30,640
you will get the word list for all of them.

120
00:08:30,640 --> 00:08:34,300
So no matter whether it is list or listed or listing,

121
00:08:34,300 --> 00:08:38,330
it still gives you the stem of a word as list.

122
00:08:38,330 --> 00:08:43,480
This is advantageous because you can now count the frequency of

123
00:08:43,480 --> 00:08:49,780
list as the list word occurring itself or in any of its derivation forms,

124
00:08:49,780 --> 00:08:52,965
any of its morphological variants.

125
00:08:52,965 --> 00:08:54,830
Do you want to do it that way?

126
00:08:54,830 --> 00:08:56,495
That's a call that you have to make.

127
00:08:56,495 --> 00:08:59,945
You really want to distinguish list and listing,

128
00:08:59,945 --> 00:09:02,310
which has slightly different meaning.

129
00:09:02,310 --> 00:09:06,185
So you may probably not want to do that but you may want to do

130
00:09:06,185 --> 00:09:11,030
list and lists to be merged together and just count as one word.

131
00:09:11,030 --> 00:09:13,212
So it is a matter of choice here.

132
00:09:13,212 --> 00:09:16,545
Porter stemmer has a particular algorithm to do it

133
00:09:16,545 --> 00:09:21,835
and it just makes all of these words the same word, list.

134
00:09:21,835 --> 00:09:26,331
A slight variant of stemming is lemmatization.

135
00:09:26,331 --> 00:09:30,280
Lemmatization is where you want

136
00:09:30,280 --> 00:09:34,790
to have the words that come out to be actually meaningful.

137
00:09:34,790 --> 00:09:36,775
Let's take an example.

138
00:09:36,775 --> 00:09:45,380
NLTK has a corpus of the universal declaration of human rights as one of its corpus.

139
00:09:45,380 --> 00:09:47,555
So if you say nltk.corpus.udhr,

140
00:09:47,555 --> 00:09:54,450
that is the Universal Declaration of Human Rights, dot words,

141
00:09:54,450 --> 00:09:56,260
and then they are end quoted with English Latin,

142
00:09:56,260 --> 00:10:02,390
this will give you all the entire declaration as a variable udhr.

143
00:10:02,390 --> 00:10:07,530
So, if you just print out the first 20 words,

144
00:10:07,530 --> 00:10:10,091
you'll see that Universal Declaration of Human Rights and there

145
00:10:10,091 --> 00:10:12,630
is a preamble and then it starts as whereas

146
00:10:12,630 --> 00:10:15,210
recognition of the inherent dignity and of

147
00:10:15,210 --> 00:10:18,475
the equal and inalienable rights of people and so on.

148
00:10:18,475 --> 00:10:20,745
So it continues that way.

149
00:10:20,745 --> 00:10:29,360
Now, if you use the Porter stemmer on these words and get the stemmed version,

150
00:10:29,360 --> 00:10:33,900
you'll see that it takes out these common suffixes.

151
00:10:33,900 --> 00:10:37,016
So universal became universe without really an e at

152
00:10:37,016 --> 00:10:40,785
the end and declaration became declar,

153
00:10:40,785 --> 00:10:43,890
and of is of and human right is the same,

154
00:10:43,890 --> 00:10:47,180
rights became right, and so on.

155
00:10:47,180 --> 00:10:52,120
But now you see that univers and declar are not really valid words.

156
00:10:52,120 --> 00:10:55,541
So lemmatization would do that stemming,

157
00:10:55,541 --> 00:11:00,045
but really keep the resulting tense to be valid words.

158
00:11:00,045 --> 00:11:03,565
It is sometimes useful because you want to somehow normalize it,

159
00:11:03,565 --> 00:11:07,445
but normalize it to something that is also meaningful.

160
00:11:07,445 --> 00:11:13,630
So we could use something like a wordnet lemmatizer that NLTK provides.

161
00:11:13,630 --> 00:11:17,581
So you have nltk.WordNetLemmatizer and then if you lemmatize

162
00:11:17,581 --> 00:11:22,530
the word from the set that you've been looking so far,

163
00:11:22,530 --> 00:11:26,345
what you get is universal declaration of human rights preamble,

164
00:11:26,345 --> 00:11:28,200
whereas recognition of the inherent dignity,

165
00:11:28,200 --> 00:11:31,755
so basically all these words are valid.

166
00:11:31,755 --> 00:11:34,530
How do you know that lemmatizer has worked?

167
00:11:34,530 --> 00:11:39,070
If you look at the first string up there and then the last string down here,

168
00:11:39,070 --> 00:11:41,815
rights has changed to right.

169
00:11:41,815 --> 00:11:44,985
So it has lemmatized it.

170
00:11:44,985 --> 00:11:48,786
But you will also notice that the fifth word here,

171
00:11:48,786 --> 00:11:56,040
universal declaration of human rights is not lemmatized because that is with a capital R,

172
00:11:56,040 --> 00:12:00,030
it's a different word that was not lemmatized to right.

173
00:12:00,030 --> 00:12:02,135
But if you had them in lower case,

174
00:12:02,135 --> 00:12:04,260
then the rights would become right again, okay?

175
00:12:04,260 --> 00:12:09,030
So there are rules of why something was lemmatized and something was kept as is.

176
00:12:09,030 --> 00:12:13,200
Once we have handled stemming and lemmatization,

177
00:12:13,200 --> 00:12:17,585
let us take a step back and look at the tokens themselves.

178
00:12:17,585 --> 00:12:19,755
The task of tokenizing something.

179
00:12:19,755 --> 00:12:24,463
So recall that we looked at how to split

180
00:12:24,463 --> 00:12:30,250
a sentence into words and tokens and we said we could just split on space.

181
00:12:30,250 --> 00:12:32,952
Right? So if you take a text string like this text11 is,

182
00:12:32,952 --> 00:12:35,883
"Children shouldn't drink a sugary drink before bed. "

183
00:12:35,883 --> 00:12:39,300
And you split on space,

184
00:12:39,300 --> 00:12:41,490
you'll get these words.

185
00:12:41,490 --> 00:12:43,905
Children shouldn't as one word,

186
00:12:43,905 --> 00:12:46,080
drink a sugary drink before bed,

187
00:12:46,080 --> 00:12:49,400
but unfortunately, you have a full stop that goes with bed.

188
00:12:49,400 --> 00:12:51,250
So it's bed full stop.

189
00:12:51,250 --> 00:12:52,860
Okay. So you got, one, two, three, four, five, six, seven,

190
00:12:52,860 --> 00:12:59,370
eight – you got eight words out of this sentence.

191
00:12:59,370 --> 00:13:05,450
But you can already see that it is not really doing a good job because,

192
00:13:05,450 --> 00:13:09,385
for example, it is keeping full stop with the word.

193
00:13:09,385 --> 00:13:14,110
So you could use the NLTK's inherent or inbuilt tokenizer,

194
00:13:14,110 --> 00:13:19,060
the way to call it would be nltk.word_tokenize and can pass the string

195
00:13:19,060 --> 00:13:25,240
there and you'll get this nice tokenized sentence.

196
00:13:25,240 --> 00:13:26,710
And in fact, it differs in two places.

197
00:13:26,710 --> 00:13:30,555
Not only is full stop taken away as a separate token,

198
00:13:30,555 --> 00:13:38,989
but you will notice that shouldn't became should and this "n't" that stands for "not",

199
00:13:38,989 --> 00:13:45,245
and that is important in quite a few NLP task because you want to know negation here.

200
00:13:45,245 --> 00:13:52,695
And the way you would do it would be to look for tokens that are a representation of not.

201
00:13:52,695 --> 00:13:56,130
So "n't" is one such representation.

202
00:13:56,130 --> 00:14:00,340
But now you know that this particular sentence does not really have

203
00:14:00,340 --> 00:14:06,685
eight tokens but 10 of them because you've got "n't" and full stop has two new tokens.

204
00:14:06,685 --> 00:14:09,415
So we talked about tokenizing

205
00:14:09,415 --> 00:14:13,800
a particular sentence and the fact that these punctuation marks have to be separated,

206
00:14:13,800 --> 00:14:21,260
there are some unique words like n apostrophe t that should also be separated and so on.

207
00:14:21,260 --> 00:14:25,360
But there is even more fundamental question of,

208
00:14:25,360 --> 00:14:29,610
what is a sentence and how do you know sentence boundaries?

209
00:14:29,610 --> 00:14:32,740
And the reason why that is important is because you

210
00:14:32,740 --> 00:14:35,855
want to split sentences from a long text sentence, right?

211
00:14:35,855 --> 00:14:40,370
So suppose this example of text12 is, this is the first sentence.

212
00:14:40,370 --> 00:14:42,730
A gallon of milk in the U.S. costs $2 99.

213
00:14:42,730 --> 00:14:44,650
And is this a third sentence, a question mark.

214
00:14:44,650 --> 00:14:48,170
And yes, it is with an exclamation.

215
00:14:48,170 --> 00:14:51,595
So already, you know that a sentence can

216
00:14:51,595 --> 00:14:56,880
end with a full stop or a question mark or an exclamation mark and so on.

217
00:14:56,880 --> 00:15:00,880
But, not all full stops and sentences.

218
00:15:00,880 --> 00:15:02,915
So for example, U dot S dot,

219
00:15:02,915 --> 00:15:06,415
that stands for US is just one word,

220
00:15:06,415 --> 00:15:07,610
has two full stops,

221
00:15:07,610 --> 00:15:10,510
but neither of them end the sentence.

222
00:15:10,510 --> 00:15:12,925
The same thing with $2.99.

223
00:15:12,925 --> 00:15:18,895
That full stop is an indicator of a number but not end of a sentence.

224
00:15:18,895 --> 00:15:22,410
We could use NLTK's inbuilt sentence splitter here and

225
00:15:22,410 --> 00:15:28,240
if you say something like nltk.sent_tokenize instead of word tokenize,

226
00:15:28,240 --> 00:15:29,868
sent tokenize and pass the string,

227
00:15:29,868 --> 00:15:32,695
it will give you sentences.

228
00:15:32,695 --> 00:15:35,435
If you count the number of sentences in this particular case,

229
00:15:35,435 --> 00:15:36,605
we should have four.

230
00:15:36,605 --> 00:15:40,900
Yey! We got four. The sentences themselves are exactly what we expect.

231
00:15:40,900 --> 00:15:42,925
This is the first sentence, is the first one.

232
00:15:42,925 --> 00:15:45,319
A gallon of milk in the US cost $2.99,

233
00:15:45,319 --> 00:15:46,805
that's the second one.

234
00:15:46,805 --> 00:15:48,285
Is this the third sentence?

235
00:15:48,285 --> 00:15:51,840
That's the third one. And yes it is, is the fourth one.

236
00:15:51,840 --> 00:15:54,545
So, what did you learn here?

237
00:15:54,545 --> 00:15:59,050
NLTK is a widely used toolkit for text and natural language processing.

238
00:15:59,050 --> 00:16:02,960
It has quite a few tools and very handy tools

239
00:16:02,960 --> 00:16:07,474
to tokenize and split a sentence and then go from there,

240
00:16:07,474 --> 00:16:09,960
lemmatize and stem and so on.

241
00:16:09,960 --> 00:16:12,950
It gives access to many text corpora as well.

242
00:16:12,950 --> 00:16:16,700
And these tasks of sentence splitting and tokenization and

243
00:16:16,700 --> 00:16:22,220
lemmatization are quite important preprocessing tasks and they are non-trivial.

244
00:16:22,220 --> 00:16:24,830
So you cannot really write a regular expression

245
00:16:24,830 --> 00:16:29,280
in a trivial fashion and expect it to work well.

246
00:16:29,280 --> 00:16:31,160
And NLTK gives you access to

247
00:16:31,160 --> 00:16:36,840
the best algorithms or at least the most suitable algorithms for these tasks.