1
00:00:00,000 --> 00:00:05,557
Let's now look at a particular model for
realizing lexicalized PCFGs. And the model

2
00:00:05,557 --> 00:00:11,528
we're going to look at is the model of
Eugene Charniak from Charniak 1997. This

3
00:00:11,528 --> 00:00:16,136
isn't the most recent model, but I'm
choosing it because it's the simplest and

4
00:00:16,136 --> 00:00:21,359
most straightforward way of building a
lexicalized PCFG. And so hopefully, it's

5
00:00:21,359 --> 00:00:26,657
easy for you guys to get a sense of how it
works. So just to give a bit of context

6
00:00:26,657 --> 00:00:31,890
for what I'm explaining. When you're
actually parsing in Charniak's parsing

7
00:00:31,890 --> 00:00:36,719
model, the way he's doing the parsing is
bottom-up in a way somewhat similar to the

8
00:00:36,719 --> 00:00:42,536
way that we did CKY parsing in an earlier
segment. But, just as with the plain

9
00:00:42,536 --> 00:00:47,094
vanilla PCFG, the probabilistic
conditioning is top-down, so you've got

10
00:00:47,094 --> 00:00:51,618
the probability of a right-hand side given
a left-hand side, i.e. the

11
00:00:51,618 --> 00:00:55,494
probability of stuff below given the stuff
above. And in this segment I'm basically

12
00:00:55,494 --> 00:00:59,655
just gonna show you what the probability
distributions are, rather than actually

13
00:00:59,655 --> 00:01:04,667
concretely going through the parsing
algorithm, but for the actual parsing algorithm

14
00:01:04,667 --> 00:01:09,680
you're using these probabilities and
applying them working upwards through the tree.

15
00:01:10,100 --> 00:01:15,408
So this is the idea of how Charniak's
algorithm works. And for this example,

16
00:01:15,408 --> 00:01:20,301
what I'm going to assume is that this is
our starting point, where we're already

17
00:01:20,301 --> 00:01:26,656
partway through building a tree. So we
have an S node, and we know the headword

18
00:01:26,662 --> 00:01:31,941
of the whole sentence is going to be rose.
And that sentence is rewriting as a noun

19
00:01:31,941 --> 00:01:38,299
phrase and a verb phrase. Well, since the
verb phrase is the head of the sentence,

20
00:01:38,299 --> 00:01:44,049
we just know automatically that its
headword is also going to be rose. And the

21
00:01:44,049 --> 00:01:49,727
point at which we are up to is, we haven't
yet decided how to expand this noun

22
00:01:49,727 --> 00:01:56,980
phrase. And we don't know what the head of
this noun phrase is. So, in Charniak's model,

23
00:01:56,980 --> 00:02:03,293
there are two probability distributions
that are used to expand out a sentence and

24
00:02:03,293 --> 00:02:10,504
we'll see each of them in turn. So, the
first one is, gee, we have to choose some

25
00:02:10,504 --> 00:02:16,333
headword for this noun phrase. And for
choosing that headword, we're going to

26
00:02:16,333 --> 00:02:24,680
condition on several things. We're going to
condition on the headword of the parent,

27
00:02:24,680 --> 00:02:33,215
the category, which is noun phrase, and
the parent category, which here is S.

28
00:02:33,215 --> 00:02:38,612
And so the probability distribution we're
going to use is this one. So we're going

29
00:02:38,612 --> 00:02:45,025
to have a probability over different
headword choices given these three

30
00:02:45,025 --> 00:02:53,660
conditioning things: the category, the
parent category, and the parent head word.

31
00:02:53,660 --> 00:03:03,795
And so what we are asking for a noun
phrase, which is in this context of being

32
00:03:03,795 --> 00:03:14,097
under an S node and having a head of the
whole S be rose, what are likely nouns to

33
00:03:14,097 --> 00:03:19,885
choose as the head of the noun phrase. And
so you might think of something like, the

34
00:03:19,885 --> 00:03:25,045
balloon rose and say balloon or
temperature, temperature rose. Or you might

35
00:03:25,045 --> 00:03:30,501
think tempers or something like this. But
since actually here we're dealing with the

36
00:03:30,501 --> 00:03:35,353
financial newspaper, the Wall Street
Journal, what it's actually likely to be,

37
00:03:35,353 --> 00:03:42,154
and is in this example, is that profits
rose. Okay so now we have this noun phrase

38
00:03:42,154 --> 00:03:48,128
here, which is a noun phrase, headed by
profits. But we don't actually know how

39
00:03:48,128 --> 00:03:53,255
it expands and so now we're going to have a
second probability distribution at work

40
00:03:53,255 --> 00:03:59,725
and so we are going to want to have a
rewrite rule that's working out how this

41
00:03:59,725 --> 00:04:06,697
expands and what we're going to condition
it on is the headword profits, the

42
00:04:06,697 --> 00:04:12,278
category noun phrase, and the parent
category which is S. And so that's then

43
00:04:12,278 --> 00:04:18,063
going to give us this second probability
distribution here. What is the probability

44
00:04:18,063 --> 00:04:23,737
of a rule that expands us down one level
in the tree, given these three

45
00:04:23,737 --> 00:04:29,883
conditioning things? The category, the
parent category, and the headword. Okay so

46
00:04:29,883 --> 00:04:37,819
we're gonna choose a rule. And so that
rule could be just NP goes to a noun,

47
00:04:37,819 --> 00:04:43,551
a plural noun and it would have some
probability or it could be something else.

48
00:04:43,551 --> 00:04:48,821
And so for this example here the actual
rule that's chosen is to generate an

49
00:04:48,821 --> 00:04:55,235
adjective and a plural noun. And so that's
given by this probability here. Well, once

50
00:04:55,235 --> 00:05:00,793
we've generated an adjective and a plural
noun, we then have the notion of a head of

51
00:05:00,793 --> 00:05:05,565
a phrase. So we know deterministically
that if this, if this is a noun phrase

52
00:05:05,565 --> 00:05:11,013
headed by profits well this plural noun
must be profits because it is the head of

53
00:05:11,013 --> 00:05:16,395
the noun phrase. And since actually we're
down to the level of a part of speech tag

54
00:05:16,395 --> 00:05:22,029
here, that means actually we have the word
profits down at the bottom of the tree

55
00:05:22,029 --> 00:05:28,475
here. Okay, at this point we've
completed this element of the expansion

56
00:05:28,475 --> 00:05:33,831
according to these two probability
distributions. And at that point, we're

57
00:05:33,831 --> 00:05:38,967
just going to keep on going. So, we're
going to say, well here is a, is a

58
00:05:38,967 --> 00:05:44,837
category, it's actually a non-terminal
category. But we're going to expand that

59
00:05:44,837 --> 00:05:50,340
by working out what its headword is. So,
we're going to use this probability

60
00:05:50,340 --> 00:05:56,471
distribution and we're going to choose a
word here as the headword corporate. And then

61
00:05:56,471 --> 00:06:01,604
by that point because we're down to a
non-terminal, we'll know that that's the

62
00:06:01,604 --> 00:06:08,418
word in the sentence. And then over here,
the VP, headed by rose, it's going to have

63
00:06:08,418 --> 00:06:15,186
to be expanded by making use of this
probability distribution so we're going to

64
00:06:15,420 --> 00:06:23,037
expand it, and choose some expansion
phrase. Maybe it'll go to a past tense

65
00:06:23,037 --> 00:06:28,701
verb and a prepositional phrase. And then
we'll start expanding those down by

66
00:06:28,701 --> 00:06:33,080
choosing headwords and how to expand them.
We'll get the rose here automatically

67
00:06:33,080 --> 00:06:37,913
again of course, which will then go to the
word rose. So, that then we choose a head

68
00:06:37,913 --> 00:06:43,290
word over here and then start expanding
it. So, we're pretty much done. So the

69
00:06:43,290 --> 00:06:48,045
only thing we really need now is a way to
get started. And so for that we just need

70
00:06:48,045 --> 00:06:53,659
a slightly varied probability distribution
at the beginning. So that we're going to

71
00:06:53,659 --> 00:06:58,339
say that we have a root category at the
top, and it needs a headword. So at that

72
00:06:58,339 --> 00:07:05,257
point we're going to have a probability of
a headword given that the category equals

73
00:07:05,257 --> 00:07:12,736
root. And then the other things, these
don't exist. So this is just the chance of

74
00:07:12,736 --> 00:07:18,215
different words being the head of the
whole sentence. And so then once we've

75
00:07:18,215 --> 00:07:24,726
done that, we can then start expanding
downwards from there. So this will be a, a

76
00:07:24,726 --> 00:07:30,992
root head, headed by rose, and then we'll
look for a rule to expand that. So we're

77
00:07:30,992 --> 00:07:36,698
now looking for some expansion, which will
here just be to a S under this. So at

78
00:07:36,698 --> 00:07:43,954
this point we again don't yet have a
parent category but we've now got a

79
00:07:43,954 --> 00:07:49,087
category and a head. So you need a couple
of special probability distributions to

80
00:07:49,087 --> 00:07:53,664
just get you started at the beginning and
then you do the basic recursion I've

81
00:07:53,664 --> 00:08:01,243
talked of here. So what we achieve by
putting these words into the grammar like

82
00:08:01,243 --> 00:08:06,159
this. So here are some statistics that
show this, that go in a bit more detail

83
00:08:06,159 --> 00:08:11,895
what I talked about previously. So here we
have different expansions of the verb

84
00:08:11,895 --> 00:08:16,789
phrase rule, and here we have a choice of
different verbs for a head. And this just

85
00:08:16,789 --> 00:08:22,555
illustrates more systematically how the
chance of different expansions for the

86
00:08:22,555 --> 00:08:28,653
verb phrase vary enormously depending on
which verb you're choosing so if you have

87
00:08:28,653 --> 00:08:34,310
a verb like come, you find out that
getting a PP after the verb is enormously,

88
00:08:34,310 --> 00:08:40,114
enormously common. That happens about
one-third of the time and having just a

89
00:08:40,114 --> 00:08:44,940
verb in the verb phrase happens about ten
percent of the time where many other kinds

90
00:08:44,940 --> 00:08:51,225
of complements like S and NP complements
are really, really rare with the verb

91
00:08:51,225 --> 00:08:57,090
come. But that's then exactly-- and also
tran--, with a noun phrase after it. But

92
00:08:57,090 --> 00:09:02,062
for other verbs the facts are very
different. So for the verb take, well, take

93
00:09:02,062 --> 00:09:07,834
normally takes an object, He took a nap,
he took a book, he took a ticket,

94
00:09:07,834 --> 00:09:13,887
anything like that. So about a third of
the time, you get just the VP goes to V NP,

95
00:09:13,887 --> 00:09:21,056
though of course you can get other things
as well. If we then move on to think,

96
00:09:21,384 --> 00:09:26,728
think is a, sort of sentential complement verb
that you're saying what you thought, and

97
00:09:26,728 --> 00:09:31,657
in fact nearly always with think, almost
three quarters of the time, you're getting

98
00:09:31,657 --> 00:09:36,345
an SBAR complement, that's something
like, he thinks that she is dishonest, he

99
00:09:36,345 --> 00:09:40,853
thinks she is dishonest. Either of those
is an SBAR complement regardless

100
00:09:40,853 --> 00:09:45,782
whether there's an overt that or not. And
in the final example illustrated here is

101
00:09:45,782 --> 00:09:49,990
the verb want, which also takes
complements but it takes infinitive S

102
00:09:49,990 --> 00:09:54,738
complements. I want to go to the store,
and so about seventy percent of the time,

103
00:09:54,738 --> 00:09:58,973
you get that. And so, essentially you
notice that they're just extremely

104
00:09:58,973 --> 00:10:03,557
different probabilities of different
expansions to the verb phrase, dependent

105
00:10:03,557 --> 00:10:08,559
on knowing what the head verb is. And this
is precisely the kind of information that

106
00:10:08,559 --> 00:10:13,453
can be captured in the Charniak parse
lexicalized parsing model. And one other

107
00:10:13,453 --> 00:10:19,112
final thing to note here, is that these
are what are sometimes referred to as

108
00:10:19,112 --> 00:10:25,604
mono-lexical probabilities. So we're
looking at the expansion of categories in

109
00:10:25,604 --> 00:10:32,715
terms of the categories and knowing just
one lexical item, the head verb. That's in

110
00:10:32,715 --> 00:10:39,652
slight contrast to the other main way in
which lexicalized PCFGs are useful, and

111
00:10:39,652 --> 00:10:45,095
that's for predicting dependencies in
things like prepositional phrases. So

112
00:10:45,095 --> 00:10:54,200
there we have bilexical probabilities. So
that was, when we had examples like, go

113
00:10:54,200 --> 00:11:02,541
into or man with, we're then deciding how
likely the connection is between the

114
00:11:02,541 --> 00:11:10,724
preposition and a noun or a verb. So these
then involve two words at a time. So the

115
00:11:10,724 --> 00:11:18,420
Charniak model also has bilexical
probabilities. So here are the chances of

116
00:11:18,420 --> 00:11:27,426
choosing the head noun of a noun phrase,
a plural noun phrase given some amount of

117
00:11:27,426 --> 00:11:32,919
information, which might include
information about the noun phrase, but

118
00:11:32,919 --> 00:11:39,274
also information about what its parent
category is and what the headword of the,

119
00:11:39,274 --> 00:11:45,473
the sentence, whole sentence is. And so
what you find in the Wall Street Journal,

120
00:11:45,473 --> 00:11:51,182
this isn't typical of all English, is that
if you have a plural noun a bit over one

121
00:11:51,182 --> 00:11:58,749
percent of the noun it's prices. But if you
know that it's the subject, it's inside a noun phrase

122
00:11:58,749 --> 00:12:04,564
under an S node, i.e. it's the subject of the
sentence, well then the chances of being

123
00:12:04,564 --> 00:12:10,288
prices become about two and a half
percent. But what really makes a big

124
00:12:10,288 --> 00:12:16,340
difference is that if you know that the
verb is fell. Well. That's information

125
00:12:16,340 --> 00:12:22,297
that tells you it's the kind of verb that
could easily go with prices. And so then

126
00:12:22,297 --> 00:12:27,344
the probability becomes far higher again,
so now it's up to an almost to fifteen

127
00:12:27,344 --> 00:12:32,208
percent chance, a one in seven chance,
that the head of the noun phrase is

128
00:12:32,208 --> 00:12:36,719
gonna be prices. And again we are
capturing a lot more of this

129
00:12:36,719 --> 00:12:41,748
probabilistic information. But you might
be wondering now, gee, can we really

130
00:12:41,748 --> 00:12:46,350
estimate these probabilities? And the
answer is that in general, you can't

131
00:12:46,350 --> 00:12:51,392
estimate these probabilities. That the
probabilities that you'd like to estimate

132
00:12:51,392 --> 00:12:56,498
become far too sparse. And I'll illustrate
that on the next slide. But the way that

133
00:12:56,498 --> 00:13:01,289
it's dealt with in Charniak's model, is by
having a complicated scheme of doing

134
00:13:01,289 --> 00:13:06,684
linear interpolation between different
models that are more or less precise. So,

135
00:13:06,684 --> 00:13:12,110
this should be reminiscent of what you saw
for language models, in the second

136
00:13:12,110 --> 00:13:17,774
week of the class. So we want to estimate
this probability distribution that we've

137
00:13:17,774 --> 00:13:23,110
seen before, choosing a headword based on
the parent's headword, your current

138
00:13:23,110 --> 00:13:28,316
category and the parent's category. And
the way you were doing that is by taking

139
00:13:28,316 --> 00:13:33,262
this linear interpolation of a bunch of
different distributions, that, one of

140
00:13:33,262 --> 00:13:38,403
which is just the maximum likelihood
estimation conditioning on everything. And

141
00:13:38,403 --> 00:13:43,571
then there are further distributions that
first of all leave out what the parent

142
00:13:43,571 --> 00:13:48,573
headword is, and then also leave out even
the parent category. So at this point,

143
00:13:48,573 --> 00:13:53,764
you're just choosing a headword based on
the category. So just saying it's a noun

144
00:13:53,764 --> 00:13:59,354
phrase, what's the chance of it having a
certain head. And in Charniak's model, these

145
00:13:59,354 --> 00:14:06,401
different distributions are weighted in a
deterministic way, depending on how much

146
00:14:06,401 --> 00:14:13,828
you'd expect to have seen certain kinds of
evidence. So, making use of these language

147
00:14:13,828 --> 00:14:20,299
modeling like techniques, are essential to
build these kind of lexicalized PCFGs,

148
00:14:20,299 --> 00:14:27,441
because the data just is too sparse. And
this next slide shows that. So, here

149
00:14:27,441 --> 00:14:33,501
the different distributions are being combined
together in the linear interpolation. So

150
00:14:33,501 --> 00:14:39,488
if we do the one on the right first, what
we find out is that if you've got a noun

151
00:14:39,488 --> 00:14:45,694
phrase headed by profits, and inside
it there's an adjective and you're wanting

152
00:14:45,694 --> 00:14:51,084
to ask what adjective it is. In the Wall
Street Journal, about one quarter of the

153
00:14:51,084 --> 00:14:56,823
time it's corporate profits. Okay. So
that's a very precise, fully conditioned

154
00:14:56,823 --> 00:15:03,230
estimate. Whereas when you start erasing
some of the information, you get.

155
00:15:03,230 --> 00:15:08,618
coarser, but still non-zero estimates.
I haven't actually made sure this second

156
00:15:08,618 --> 00:15:13,056
line here, this was the method that
Charniak used to get kind of coarse

157
00:15:13,056 --> 00:15:18,128
semantic classes, to try and keep some
information about parent head word before

158
00:15:18,128 --> 00:15:23,136
getting rid of it entirely. But if we just
look at--, don't really do a lot

159
00:15:23,136 --> 00:15:27,062
with that one, and look at these two. You
can see that once you're only,

160
00:15:27,062 --> 00:15:32,836
conditioning on knowing it's a noun phrase
under an S or just a noun phrase, the

161
00:15:32,836 --> 00:15:37,196
chance of it being corporate, is then
dropping by almost two orders of

162
00:15:37,196 --> 00:15:43,291
magnitude. So you're down to about half a
percent. Okay, so this is the good case in

163
00:15:43,291 --> 00:15:49,002
which you can calculate a probability from
the rich conditioning, and it helps you a

164
00:15:49,002 --> 00:15:56,128
lot. But quite commonly, that just doesn't
work for you. So if you then say, well, the

165
00:15:56,128 --> 00:16:00,767
verb of the sentence is rose, and I've got
a subject noun phrase, a noun phrase under

166
00:16:00,767 --> 00:16:08,835
an S, And what, noun should it be? Well,
in the particular sentence in the data to

167
00:16:08,835 --> 00:16:16,047
be parsed, the actual noun, was profits.
But it turns out that in the training data,

168
00:16:16,047 --> 00:16:22,233
profits never occurred as the noun heading
a noun phrase that was the subject of

169
00:16:22,233 --> 00:16:27,291
rose, despite the fact that that sounds
perfectly normal. Profits rose last

170
00:16:27,291 --> 00:16:32,491
quarter. Profits rose throughout the
economy. Any sentence like that. And so

171
00:16:32,491 --> 00:16:38,323
this, the maximum likelihood estimate, the
MLE here is just zero. And so the only way

172
00:16:38,323 --> 00:16:44,092
we're getting a non-zero probability
estimate is by using these probabilities

173
00:16:44,092 --> 00:16:50,454
where we condition by less stuff. And even
then we're getting these low probabilities

174
00:16:50,454 --> 00:16:56,302
right. So the probability of a noun being
profits in a noun phrase it's about sort

175
00:16:56,302 --> 00:17:02,138
of one 20th of a percent. But that's the best
kind of estimate that we have. And so although

176
00:17:02,138 --> 00:17:07,836
you'd like to use rich estimates as
here most of the time in doing lexicalized

177
00:17:07,836 --> 00:17:13,325
PCFG parsing, you're actually having to
fall back on rather coarser estimates

178
00:17:13,325 --> 00:17:18,290
because you can't get the conditioning
information you'd like to from the fairly

179
00:17:18,290 --> 00:17:24,324
small supervised tree banks that we have
to train on. In particular, for these

180
00:17:24,324 --> 00:17:30,991
bilexical probabilities when you're
trying to condition on two lexical items,

181
00:17:30,991 --> 00:17:36,403
the probabilities tend to have to
get backed off, just because the amount of

182
00:17:36,403 --> 00:17:42,208
information to estimate this is extremely,
extremely sparse. Because, you're in this

183
00:17:42,208 --> 00:17:47,607
space where even before you consider the
categories, that you're doing something

184
00:17:47,607 --> 00:17:52,874
that's kind of like bigram probability
estimates. And it's hard to estimate word

185
00:17:52,874 --> 00:17:57,750
bigram probabilities on only about a
million words of text. And that's all the

186
00:17:57,750 --> 00:18:04,390
text that we have in the hand-constructed
tree banks. Okay. There were some details

187
00:18:04,390 --> 00:18:09,966
about, that needed smoothing at
the end there. But I hope the main thing

188
00:18:09,966 --> 00:18:14,327
that you could take away from this segment
is how there was a fairly straightforward

189
00:18:14,327 --> 00:18:20,389
system of two probability distributions
that Charniak was able to use to realize a

190
00:18:20,389 --> 00:18:22,800
lexicalized PCFG model.
