1
00:00:01,300 --> 00:00:05,371
So let me just conclude this lecture with 
some results on the parsing problem using 

2
00:00:05,371 --> 00:00:09,029
the reranking models that I just 
described to you, in conjunction with the 

3
00:00:09,029 --> 00:00:15,909
perception. 
So in one set of experiments by myself 

4
00:00:15,909 --> 00:00:22,365
and Terry Koon, we started with a 
baseline model which was electrolyzed. 

5
00:00:22,365 --> 00:00:27,143
PCFG. 
Which, at least, at the time, was very 

6
00:00:27,143 --> 00:00:34,078
close to the state of the art in parsing. 
And that scores around 88% precision and 

7
00:00:34,078 --> 00:00:38,654
recall. 
Remember, f measure is a kind of average 

8
00:00:38,654 --> 00:00:45,230
of precision and recall, in recovering 
sub constituents within a parse true. 

9
00:00:47,930 --> 00:00:53,759
The reranked model that I just described 
scores 89.5% f-measure, which is about 

10
00:00:53,759 --> 00:01:00,284
11% relative error reduction. 
So about 11% of errors have been 

11
00:01:00,284 --> 00:01:08,256
corrected by the reranking model. 
So that's a fairly significant 

12
00:01:08,256 --> 00:01:14,018
improvement, given that these models are 
starting to reach quite high levels of 

13
00:01:14,018 --> 00:01:18,416
accuracy. 
This is actually a pretty significant 

14
00:01:18,416 --> 00:01:22,574
improvement, and indeed I had developed 
these lexicalized P, PCFG's during my 

15
00:01:22,574 --> 00:01:26,732
Ph.D thesis, and it was very, very hard 
to push these any further other than this 

16
00:01:26,732 --> 00:01:34,337
88.2% measure which we see here. 
Here are some other results more recently 

17
00:01:34,337 --> 00:01:39,244
from Eugene Charniak and Mark Johnson in 
2005. 

18
00:01:39,244 --> 00:01:45,339
They employed a similar approach. 
But they had better and best lists, 

19
00:01:45,339 --> 00:01:48,858
better features, and also importantly a 
better baseline model than this model 

20
00:01:48,858 --> 00:01:54,440
I've shown you here. 
And they pushed accuracy from about 89.7% 

21
00:01:54,440 --> 00:01:58,020
to 91% accuracy. 
This is actually very, very close to the 

22
00:01:58,020 --> 00:02:02,833
state of the art in parsing performance. 
So, the reranking model, again, gives a 

23
00:02:02,833 --> 00:02:07,306
pretty significant gain and actually 
produced one of the very best results 

24
00:02:07,306 --> 00:02:14,900
we've seen on parsing. 
What I've shown you, though, in this 

25
00:02:14,900 --> 00:02:19,116
lecture, is a quite new way of thinking 
about these supervised learning problems 

26
00:02:19,116 --> 00:02:23,650
that we see in natural language 
processing. 

27
00:02:23,650 --> 00:02:28,834
This idea of global linear models defined 
through gen f and v, and finally the 

28
00:02:28,834 --> 00:02:35,199
perception algorithm as one way of 
training these parameters v. 

29
00:02:35,199 --> 00:02:40,083
They give significant improvements on 
these reranking problems, but perhaps 

30
00:02:40,083 --> 00:02:44,523
most importantly they're going to open up 
a whole new way of thinking about 

31
00:02:44,523 --> 00:02:51,350
algorithms for problems such as 
translation or tagging and parsing. 

32
00:02:51,350 --> 00:02:55,070
And we'll see how we can apply these 
models in several other contexts in the 

33
00:02:55,070 --> 00:02:59,640
final week of lectures, which is the next 
week of this class. 

