1
00:00:08,867 --> 00:00:13,843
In this video, we are going to
move on from basic NLP tasks to

2
00:00:13,843 --> 00:00:16,540
advanced NLP tasks using NLTK.

3
00:00:17,780 --> 00:00:21,550
If you recall the NLP
tasks that we look so

4
00:00:21,550 --> 00:00:26,410
far are counting words, counting
frequency of words, finding unique words,

5
00:00:26,410 --> 00:00:29,650
finding sentence boundaries,
even finding tokens in stemming.

6
00:00:31,920 --> 00:00:32,630
In this video,

7
00:00:32,630 --> 00:00:37,830
we will talk about part of speech tagging
and parsing the sentence structure.

8
00:00:37,830 --> 00:00:40,570
But other NLP tasks like
semantic role labeling and

9
00:00:40,570 --> 00:00:43,280
named entity recognition,
that we'll cover later on.

10
00:00:46,110 --> 00:00:49,260
Let's start with part-of-speech tagging,
or POS tagging.

11
00:00:50,540 --> 00:00:54,860
Recall from your high school grammar
that part-of-speech are these

12
00:00:54,860 --> 00:00:58,949
verb classes like nouns,
and verbs, and adjectives.

13
00:01:00,810 --> 00:01:03,620
There are many, many more tags than these.

14
00:01:03,620 --> 00:01:06,200
You can see that you have conjunctions and
cardinals.

15
00:01:06,200 --> 00:01:11,520
Cardinals are, if you have a number, then
you are to kind of assign that word class.

16
00:01:11,520 --> 00:01:12,520
You have determiner,

17
00:01:12,520 --> 00:01:17,030
you have prepositions,
you have modal words, you have nouns.

18
00:01:17,030 --> 00:01:21,780
Again, nouns could be singular nouns,
and plural nouns, and proper nouns.

19
00:01:21,780 --> 00:01:25,640
You have possessives and
pronouns again of multiple types,

20
00:01:25,640 --> 00:01:29,450
you have adverbs, symbols, and then verbs.

21
00:01:29,450 --> 00:01:32,920
And verbs themselves are multiple classes,
you have verbs and

22
00:01:32,920 --> 00:01:35,530
gerunds and past tense verbs and so on.

23
00:01:37,412 --> 00:01:39,130
How to get them?

24
00:01:39,130 --> 00:01:45,000
Recall that in NLTK, you need to import
NLTK first, and then you can get more

25
00:01:45,000 --> 00:01:50,249
information about these word classes,
by issuing a help command.

26
00:01:50,249 --> 00:01:55,500
So nltk.help you've been tagset,
this comes from upenn tags.

27
00:01:55,500 --> 00:02:01,270
So you, if you say upenn_tagset,
and then give the tag there,

28
00:02:01,270 --> 00:02:04,810
MD, it'll tell you that MD stands for

29
00:02:04,810 --> 00:02:10,350
modal auxiliary, and
these are the words that are modal words.

30
00:02:10,350 --> 00:02:13,270
Can, cannot, could, couldn't, may,

31
00:02:13,270 --> 00:02:17,210
might, ought, shall, should,
shouldn't, will, would, and so on.

32
00:02:18,830 --> 00:02:22,500
So how do you do POS tagging with NLTK?

33
00:02:22,500 --> 00:02:24,946
Recall that you have to split
a sentence into words, okay.

34
00:02:26,100 --> 00:02:28,800
So this example is something
that we have seen.

35
00:02:28,800 --> 00:02:32,290
Children shouldn't drink
a sugary drink before bed.

36
00:02:32,290 --> 00:02:38,500
But if I split it into words, so
you could use a word_tokenize

37
00:02:38,500 --> 00:02:43,810
from an NLTK package that text13
then becomes this individual word.

38
00:02:43,810 --> 00:02:46,270
Recall that there were ten words here,

39
00:02:46,270 --> 00:02:50,280
children shouldn't,
as in a separate token,

40
00:02:50,280 --> 00:02:54,480
and then drink a sugary drink before bed,
and the full stop is the last token.

41
00:02:57,210 --> 00:03:00,775
Then, if you run the post tag there,

42
00:03:00,775 --> 00:03:06,870
by using this command nltk.pos_tag
on this tokenized form,

43
00:03:06,870 --> 00:03:11,930
you'll get the tags, so
children is a plural noun,

44
00:03:11,930 --> 00:03:17,105
should is a model word,
n't is tagged as an end verb,

45
00:03:17,105 --> 00:03:23,779
drink is a verb, a is a determiner,
sugary is an adjective and so on.

46
00:03:23,779 --> 00:03:29,540
So that is how you'll
get the part of speech.

47
00:03:29,540 --> 00:03:34,729
Now, why do you want part of speech,
because if you are trying to

48
00:03:34,729 --> 00:03:40,692
club all nouns together because they
all have some sort of a similar form or

49
00:03:40,692 --> 00:03:46,749
all modulus together then using the part
of speech tag gives you one class or

50
00:03:46,749 --> 00:03:49,920
one cluster for all of these words, and

51
00:03:49,920 --> 00:03:54,180
then you don't have to
address them individually.

52
00:03:54,180 --> 00:03:56,520
So when your doing some
feature engineering or

53
00:03:56,520 --> 00:03:59,400
feature extraction that is very useful.

54
00:03:59,400 --> 00:04:03,680
We'll talk about features and
how to use it next week.

55
00:04:04,930 --> 00:04:07,570
Now, that we know how to do POS tagging,

56
00:04:07,570 --> 00:04:11,120
we should talk about
ambiguity in POS tagging.

57
00:04:12,630 --> 00:04:15,970
In fact,
ambiguity is very common in English.

58
00:04:15,970 --> 00:04:20,210
Let's take this example,
visiting aunts can be a nuisance.

59
00:04:20,210 --> 00:04:21,390
What do you think it means?

60
00:04:22,610 --> 00:04:27,510
Does it mean visiting aunts
can be a nuisance, and

61
00:04:27,510 --> 00:04:34,135
this going to visit aunt can be nuisance
or aunts who are coming in are nuisance.

62
00:04:35,650 --> 00:04:40,690
So if you tokenize the word with
visiting aunts can be a nuisance and

63
00:04:40,690 --> 00:04:45,240
do a pos_tag on that you'll get one tag.

64
00:04:45,240 --> 00:04:48,830
So when you say visiting is a verb, and

65
00:04:48,830 --> 00:04:53,460
that aunts is a noun, a plural noun.

66
00:04:54,640 --> 00:04:59,250
This representation shows that you
are doing the act of visiting.

67
00:04:59,250 --> 00:05:04,300
So visiting is the verb for the aunt,
you are going to visit your aunt.

68
00:05:06,060 --> 00:05:10,150
The alternate POS tag would be
visiting as an adjective for aunt.

69
00:05:11,270 --> 00:05:15,260
So that would mean that the aunts
are the ones who are visiting you.

70
00:05:16,310 --> 00:05:20,040
So you can see that this
particular sentence is ambiguous,

71
00:05:20,040 --> 00:05:20,780
the way it is written.

72
00:05:22,110 --> 00:05:25,245
POS tagging is not, and if fact,

73
00:05:25,245 --> 00:05:30,170
NLTK gives the first version and
not the second.

74
00:05:30,170 --> 00:05:33,770
And that is because you don't
really have a way to say,

75
00:05:33,770 --> 00:05:36,280
give me all possible variance.

76
00:05:36,280 --> 00:05:41,348
And sometimes the ones that take
precedence are the ones where you mostly

77
00:05:41,348 --> 00:05:46,349
likely see the word visiting in
the context of usages in a large corpus.

78
00:05:46,349 --> 00:05:50,429
So visiting is more often
used as a continuous or

79
00:05:50,429 --> 00:05:55,640
a form of a verb rather than
using it as an adjective.

80
00:05:56,650 --> 00:06:01,280
So the probability of finding visiting
as an adjective is lower and so

81
00:06:01,280 --> 00:06:07,360
the first word winds up, but this is a
valid alternate representation of POS tag.

82
00:06:08,800 --> 00:06:11,760
Okay, so that was POS tagging.

83
00:06:13,090 --> 00:06:15,160
Once, we have done some POS tagging,

84
00:06:15,160 --> 00:06:18,580
we can look at parsing
the sentence structure itself.

85
00:06:18,580 --> 00:06:19,310
Okay, now,

86
00:06:19,310 --> 00:06:24,940
that we have looked at how you would find
out the parts of speech of a sentence.

87
00:06:24,940 --> 00:06:27,740
Let's look at the sentence
structure itself, and

88
00:06:27,740 --> 00:06:29,100
parsing of the sentence structure.

89
00:06:30,480 --> 00:06:35,050
Making sense of sentences is easy if
you follow a well defined grammatic

90
00:06:35,050 --> 00:06:36,630
structure, okay?

91
00:06:36,630 --> 00:06:40,020
So when you have a sentence like this,
Alice loves Bob,

92
00:06:41,050 --> 00:06:46,780
you want to know which is the noun and
which is the verb and so on.

93
00:06:46,780 --> 00:06:51,090
But we also want to know how
are they related in the sentence?

94
00:06:51,090 --> 00:06:54,102
How did they come together
to form the sentence?

95
00:06:54,102 --> 00:06:59,720
So in NLTK, you'll use
nltk.word_tokenize Alice loves Bob.

96
00:06:59,720 --> 00:07:04,065
That'll give you these three words in
the sentence, right, Alice loves Bob.

97
00:07:04,065 --> 00:07:06,590
Those are the three words, but

98
00:07:06,590 --> 00:07:11,440
the sentence itself is
constituted of two things.

99
00:07:11,440 --> 00:07:15,999
You have noun phrase and verb phrase.

100
00:07:15,999 --> 00:07:20,010
The word phrase itself can be
a word followed by a noun phrase.

101
00:07:21,080 --> 00:07:22,950
That is the grammar extraction in English.

102
00:07:25,230 --> 00:07:29,270
And for this particular sentence this
is the grammar structure that is used.

103
00:07:29,270 --> 00:07:33,780
Now, that we have that, noun trace itself
can be a noun, so Alice can be a noun and

104
00:07:33,780 --> 00:07:38,460
then Bob is the other noun and
there is one verb that is loves.

105
00:07:39,920 --> 00:07:45,071
So by doing this structure and
writing out this grammar where

106
00:07:45,071 --> 00:07:50,323
you have these grammar rules on
the right saying S gives NP VP,

107
00:07:50,323 --> 00:07:55,981
VP is V and NP, and then Alice and
Bob can be NP and loves would be a V.

108
00:07:55,981 --> 00:08:00,250
You have created what is
called a context free grammar.

109
00:08:01,590 --> 00:08:05,320
And then you can use
NLTK's contextual grammar

110
00:08:05,320 --> 00:08:08,950
input statements like nltk.CFG.fromstring,

111
00:08:08,950 --> 00:08:12,590
and then you write the string out,
that will give you the grammar.

112
00:08:13,880 --> 00:08:16,390
And you can use this grammar
to parse the sentence.

113
00:08:16,390 --> 00:08:21,860
So you can use nltk.ChartParser,
you create a parser using the grammar

114
00:08:21,860 --> 00:08:27,480
that you have defined, sort is as parser,
and then parse the sentence.

115
00:08:27,480 --> 00:08:32,620
So you parse text15, and when you
parse it, it gives you parse trees.

116
00:08:32,620 --> 00:08:35,790
And then you can print that tree,
so you'll print the tree as this.

117
00:08:35,790 --> 00:08:40,260
So the way you will parse it out,
for lack of a better word,

118
00:08:40,260 --> 00:08:44,540
is you have S, that's sentence,
that constitutes two brackets.

119
00:08:44,540 --> 00:08:49,640
The NP bracket that has Alice and
a VP bracket

120
00:08:49,640 --> 00:08:54,720
where a VP bracket has a V bracket for
loves the verb and

121
00:08:54,720 --> 00:08:59,420
an empty bracket for Bob exactly
the way we have drawn the tree.

122
00:09:00,640 --> 00:09:05,270
As with part of speech
tagging parsing also is

123
00:09:05,270 --> 00:09:10,960
ambiguous and ambiguity may exist even
if sentences are grammatically correct.

124
00:09:10,960 --> 00:09:14,980
So let's take this example,
this is a very famous example.

125
00:09:14,980 --> 00:09:16,650
I saw the man with a telescope.

126
00:09:17,800 --> 00:09:21,070
Now, did you see with the telescope?

127
00:09:21,070 --> 00:09:23,779
Or did you see the man who
was holding a telescope?

128
00:09:25,540 --> 00:09:29,650
There are two meanings, and
the meaning is with respect to

129
00:09:29,650 --> 00:09:33,510
where this preposition with
the telescope gets attached.

130
00:09:33,510 --> 00:09:36,250
So this is a typical
preposition attachment problem.

131
00:09:38,090 --> 00:09:43,260
And if you look at the grammar, you
will see where that ambiguity comes in.

132
00:09:43,260 --> 00:09:46,890
The fact that this is
grammatically correct is because

133
00:09:46,890 --> 00:09:52,060
you have sentence that is noun phrase and
verb phrase.

134
00:09:52,060 --> 00:09:56,517
But then this verb phrase can
be split into two things.

135
00:09:56,517 --> 00:09:58,970
So noun phrase is I, that is clear.

136
00:09:58,970 --> 00:10:04,160
But verb phrase could either be a verb
followed by a noun phrase as in,

137
00:10:04,160 --> 00:10:06,150
saw the man with the telescope.

138
00:10:07,880 --> 00:10:13,290
Or it could be a verb phrase followed
by a preposition phrase, where you have

139
00:10:13,290 --> 00:10:17,590
the verb phrase is, saw the man, because
verb phrase could still be verb and

140
00:10:17,590 --> 00:10:20,000
a noun phrase,
that's what the saw the man is.

141
00:10:20,000 --> 00:10:24,020
And then you have preposition phrase,
which is preposition in a noun phrase,

142
00:10:24,020 --> 00:10:26,120
so with and the telescope.

143
00:10:28,140 --> 00:10:32,820
So these are the two alternatives,
the bold red and the dotted red.

144
00:10:32,820 --> 00:10:37,150
And those are the two trees that
distinguish the two meanings of you would

145
00:10:37,150 --> 00:10:39,450
parse out the meaning from the sentence.

146
00:10:42,590 --> 00:10:44,300
We can do the same thing and

147
00:10:44,300 --> 00:10:48,452
this becomes apparent when we
use NLTK's parsing as well.

148
00:10:48,452 --> 00:10:50,775
So you first tokenize the sentence,

149
00:10:50,775 --> 00:10:54,360
nltk.word_tokenize I saw
the man with a telescope.

150
00:10:55,570 --> 00:10:57,550
And you can load a grammar, so

151
00:10:57,550 --> 00:11:03,360
if write your own file
mygrammar1.cfg that has these lines.

152
00:11:03,360 --> 00:11:05,790
That the sentence is
non-present verb phrase.

153
00:11:05,790 --> 00:11:09,230
Then verb phrase could be a verb and
a noun phrase or a verb phrase and

154
00:11:09,230 --> 00:11:10,600
a preposition phrase.

155
00:11:10,600 --> 00:11:13,980
Preposition phrase itself can be
a preposition and a noun phrase and so on.

156
00:11:15,650 --> 00:11:19,665
And then if you load this
up as a grammar using

157
00:11:19,665 --> 00:11:24,720
nltk.data.load, that
will create a grammar.

158
00:11:24,720 --> 00:11:26,950
Then you can create,
the grammar has 13 rules,

159
00:11:26,950 --> 00:11:31,270
13 productions, that is what it is called.

160
00:11:31,270 --> 00:11:35,485
And then you can create a chart parser
using this grammar, so you can say

161
00:11:35,485 --> 00:11:40,720
nltk.ChartParser{grammar1), exactly
the same way we did it a few slides back.

162
00:11:41,920 --> 00:11:45,667
And then print out all trees
from this parser from that

163
00:11:45,667 --> 00:11:50,091
I sorry the parser you will see
that there are indeed two trees.

164
00:11:50,091 --> 00:11:55,341
One, which is a noun phrase and
a verb phrase where

165
00:11:55,341 --> 00:12:01,101
verb phrase has verb phrase and
a preposition phrase.

166
00:12:01,101 --> 00:12:02,482
And another one which is a noun phrase and

167
00:12:02,482 --> 00:12:04,970
a verb phrase where verb phrase
is a verb and noun phrase.

168
00:12:04,970 --> 00:12:07,250
And the noun phrase itself
has a preposition with it.

169
00:12:08,660 --> 00:12:11,780
So these are the two parse tree structures

170
00:12:11,780 --> 00:12:14,340
that you'll get when you parse
using this context with grammar.

171
00:12:15,430 --> 00:12:21,330
Now, we gave examples of simple grammars,
and said

172
00:12:21,330 --> 00:12:26,100
we'll create a context for grammar out of
it, but we can not do that every time.

173
00:12:26,100 --> 00:12:30,510
In fact, generating the grammar and
generating grammar rules itself is

174
00:12:30,510 --> 00:12:36,010
a learning task that you could learn, and
you need a lot of training data for that.

175
00:12:36,010 --> 00:12:41,220
And a lot of manual effort and hours
have gone into creating what is known

176
00:12:41,220 --> 00:12:47,370
as a tree back, basically,
a big collection of parse trees.

177
00:12:47,370 --> 00:12:52,540
From Wall Street Journal, and in fact, if
we have access to treebank through NLTK.

178
00:12:52,540 --> 00:12:56,534
So if you say from NLTK corpus
import treebank and say

179
00:12:56,534 --> 00:13:03,100
treebank.parsed_sentences this particular
first sentence from Wall Street Journal.

180
00:13:03,100 --> 00:13:05,550
And you print out,
you will see the sentence and

181
00:13:05,550 --> 00:13:08,330
you'll realize that you
have seen it before.

182
00:13:08,330 --> 00:13:10,130
This is that Pierre Vinken sentence.

183
00:13:10,130 --> 00:13:12,930
So, Pierre Vinken, 61 years old,

184
00:13:12,930 --> 00:13:17,880
will join the board as a nonexecutive
director, November 29th.

185
00:13:17,880 --> 00:13:23,090
So that sentence has been parsed using
this structure, in the tree bank,

186
00:13:23,090 --> 00:13:24,690
and you have that available.

187
00:13:27,480 --> 00:13:33,180
Just to conclude, you see that there are
complexities in part of speech tagging and

188
00:13:33,180 --> 00:13:37,430
parsing, beyond just how to do it.

189
00:13:37,430 --> 00:13:42,640
So for example, there is this usage and
uncommon usage of words.

190
00:13:42,640 --> 00:13:47,480
An example is the old man the boat,
when you read it you

191
00:13:47,480 --> 00:13:52,750
feel that the old man is
actually a man who is old,

192
00:13:52,750 --> 00:13:56,960
but then you cannot finish the sentence,
and parse it properly.

193
00:13:56,960 --> 00:14:02,480
The current parse would make man the verb,
that is to man something.

194
00:14:04,220 --> 00:14:08,460
However, when you do a word tokenize
on the sentence do a post stack,

195
00:14:09,550 --> 00:14:13,750
you'll get man as a noun,
and old as an adjective.

196
00:14:13,750 --> 00:14:17,440
And this particular sentence
is not grammatically correct,

197
00:14:17,440 --> 00:14:19,360
there's no parse string for
this structure.

198
00:14:21,540 --> 00:14:26,530
Sometimes even well formed sentences
that have parse structures may

199
00:14:26,530 --> 00:14:32,050
still be meaningless, so there is no
semantics or meaning associated with that.

200
00:14:32,050 --> 00:14:36,650
The great example is,
colorless green ideas sleep furiously.

201
00:14:36,650 --> 00:14:40,870
When you read this you can see that
the sentence structure seems right,

202
00:14:42,220 --> 00:14:45,130
but it does not make any sense,
meaningless.

203
00:14:46,290 --> 00:14:49,040
And that is because when you do,

204
00:14:49,040 --> 00:14:53,190
when you find out that the word
tokens will be word tokenize and

205
00:14:53,190 --> 00:14:58,600
the part of speech tag on that, you will
see that it doesn't really do a good job,

206
00:14:58,600 --> 00:15:04,020
it says, colorless is a proper noun,
and then green is an adjective,

207
00:15:04,020 --> 00:15:08,110
rather than saying colorless and green are
both adjectives, there is an error there.

208
00:15:08,110 --> 00:15:11,320
But even if you remove the word colorless,
say,

209
00:15:11,320 --> 00:15:15,230
green ideas sleep furiously,
you would have the perfect post tag there.

210
00:15:15,230 --> 00:15:20,450
Green is an adjective, ideas is a noun,
sleep is a verb, and the you

211
00:15:20,450 --> 00:15:25,190
have furiously, that is an adverb, and
this particular order is perfectly fine.

212
00:15:25,190 --> 00:15:27,958
But it still doesn't make
any sense meaning wise.

213
00:15:27,958 --> 00:15:33,061
So there are many more layers of
complexity in language that parse

214
00:15:33,061 --> 00:15:37,902
trees and apart of three stacks
don't address so far, okay?

215
00:15:37,902 --> 00:15:42,904
So to conclude all of the take home
concepts here, we looked at POS tagging

216
00:15:42,904 --> 00:15:47,762
and saw how it provides insight into
the word classes, and word types.

217
00:15:47,762 --> 00:15:52,190
In the sentence, parsing the grammatical
structures helps derive meaning.

218
00:15:53,430 --> 00:15:54,731
Both tasks are difficult, and

219
00:15:54,731 --> 00:15:58,570
there is linguistic ambiguity that
increases the difficulty even more.

220
00:15:58,570 --> 00:16:02,580
And you need better models, and you can
learn those using supervised learning,

221
00:16:03,930 --> 00:16:08,080
NLTK provides access to these tools and
also has data, in terms of tree bank for

222
00:16:08,080 --> 00:16:10,410
example, that could be used for training.

223
00:16:11,410 --> 00:16:12,250
Next module,

224
00:16:12,250 --> 00:16:15,520
we're going to go into more detail
about how do you train these models,

225
00:16:15,520 --> 00:16:19,500
how do you build a supervised method,
and supervised technique, for these.

226
00:16:19,500 --> 00:16:20,750
But that's for another time.