1
00:00:00,840 --> 00:00:04,659
So the first example problem I'm going to 
use to motivate log linear models is the 

2
00:00:04,659 --> 00:00:09,762
Language Modeling Problem, which we saw 
right at the start of this class. 

3
00:00:09,762 --> 00:00:14,330
So just to recap quickly, the problem is 
as follows. 

4
00:00:14,330 --> 00:00:19,28
We define w sub i to be the ith word in a 
document, and then, our task is to 

5
00:00:19,28 --> 00:00:25,486
estimate a distribution. 
This is the conditional distribution over 

6
00:00:25,486 --> 00:00:31,720
wi given the context or history which is 
the previous i minus 1 words. 

7
00:00:31,720 --> 00:00:40,248
So here's one example, this is a passage 
taken, taken from Chomsky from the 1950s. 

8
00:00:40,248 --> 00:00:46,48
So say we have this sequence of i minus 1 
words, our task is to estimate a 

9
00:00:46,48 --> 00:00:54,400
distribution over the word that appears 
at the ith position at wi. 

10
00:00:54,400 --> 00:01:03,724
Now, of course, one model we studied 
extensively in the schools was the idea 

11
00:01:03,724 --> 00:01:11,288
of a trigram language model. 
So again, just a quick recap on this, so 

12
00:01:11,288 --> 00:01:14,554
a trigram estimate is defined as follows. 
We define this, this estimate of a 

13
00:01:14,554 --> 00:01:18,132
probability of, of a particular word, say 
model and given the previous context. 

14
00:01:18,132 --> 00:01:21,82
So, this is the word we're trying to 
predict, this is wi. 

15
00:01:21,82 --> 00:01:26,632
we defined that as a combination of three 
maximum likelihood estimates, assuming 

16
00:01:26,632 --> 00:01:31,362
that we're using smoothing. 
So we have the maximum likelihood 

17
00:01:31,362 --> 00:01:36,206
estimate of model given that the previous 
two words are any and statistical. 

18
00:01:36,206 --> 00:01:39,570
We have the maximum likelihood estimate 
of model, given just the word 

19
00:01:39,570 --> 00:01:43,398
statistical, and then finally, we have 
the so-called unigram estimate of just 

20
00:01:43,398 --> 00:01:48,316
the probability of seeing model 
conditioned on no context. 

21
00:01:48,316 --> 00:01:52,978
And of course, these maximum likelihood 
estimates are defined as ratios of 

22
00:01:52,978 --> 00:01:58,364
counts, as we've seen here. 
So if I want to estimate the probability 

23
00:01:58,364 --> 00:02:03,76
of some y given some context x, I take 
the ratio of these two counts and these 

24
00:02:03,76 --> 00:02:07,884
lambdas. 
These three values lambda 1, lambda 2, 

25
00:02:07,884 --> 00:02:14,60
lambda 3 dictate the relevant, relative 
weight of these three estimates. 

26
00:02:14,60 --> 00:02:21,750
These lambdas are positive and they sum 
to 1. 

27
00:02:21,750 --> 00:02:25,590
So that's a trigram language model, you 
should be very familiar with those at 

28
00:02:25,590 --> 00:02:30,746
this point in the class but they have 
some very clear deficiencies. 

29
00:02:30,746 --> 00:02:38,830
Okay, so these models make use of only 
bigram trigram, and unigram estimates. 

30
00:02:38,830 --> 00:02:43,48
Okay, so they basically restrict 
themselves to a very small window, 

31
00:02:43,48 --> 00:02:48,310
conditioning only on the previous two 
words in the context. 

32
00:02:48,310 --> 00:02:51,901
And if we think back to that passage, 
there could be all kinds of quote 

33
00:02:51,901 --> 00:02:56,59
features of the context, which could be 
useful in predicting the distribution 

34
00:02:56,59 --> 00:03:01,263
over the next word. 
So, for example, we might want to come up 

35
00:03:01,263 --> 00:03:05,878
with an estimate that conditions on the 
fact that the word two backs the word two 

36
00:03:05,878 --> 00:03:12,638
positions back is the word any. 
notice that we've, in some sense, skipped 

37
00:03:12,638 --> 00:03:16,780
over wi minus 1 in this case and just 
looked at the word to back. 

38
00:03:16,780 --> 00:03:20,812
We might condition on the fact that the 
previous word is an adjective, so that 

39
00:03:20,812 --> 00:03:27,28
gives us a coarser estimate which ignores 
the exact identity of the previous word. 

40
00:03:27,28 --> 00:03:30,646
but it is an estimate that we might only 
be able to make roughly reliably given 

41
00:03:30,646 --> 00:03:35,244
the amount of data we have. 
We might condition on the previous word 

42
00:03:35,244 --> 00:03:39,510
ending on a particular suffix or prefix, 
for example, ical. 

43
00:03:39,510 --> 00:03:42,60
We might condition on the author of the 
entire article. 

44
00:03:42,60 --> 00:03:45,480
That's likely to influence this 
distribution and we might condition on 

45
00:03:45,480 --> 00:03:49,560
long range features. 
So we might, for example, condition on 

46
00:03:49,560 --> 00:03:54,110
the fact that the word model the fact 
that it doesn't occur somewhere in the 

47
00:03:54,110 --> 00:03:58,645
previous context. 
Or we might condition on the fact that 

48
00:03:58,645 --> 00:04:03,710
some other word, for example grammatical, 
occurs somewhere in the previous context. 

49
00:04:03,710 --> 00:04:08,62
Notice critically, these features, and 
actually also this feature, in some 

50
00:04:08,62 --> 00:04:12,754
sense, go beyond just the two, previous 
words of context to either consider the 

51
00:04:12,754 --> 00:04:17,242
entire document or some meter information 
about the document, for example the 

52
00:04:17,242 --> 00:04:22,946
document's author. 
The point being that all of these 

53
00:04:22,946 --> 00:04:27,502
estimates could provide useful 
distribution useful information about the 

54
00:04:27,502 --> 00:04:31,786
distribution of the next word and the 
document, all of these features of the 

55
00:04:31,786 --> 00:04:37,730
context could be useful in estimating 
this distribution. 

56
00:04:37,730 --> 00:04:42,416
And of course, the trigram model ignores 
all of the information that I've shown 

57
00:04:42,416 --> 00:04:47,50
you down here. 
So lets just imagine that were trying to 

58
00:04:47,50 --> 00:04:52,730
define this estimate probability of this 
particular word model, given the previous 

59
00:04:52,730 --> 00:04:58,922
i minus 1 words in the document. 
And we want to incorporate all of these 

60
00:04:58,922 --> 00:05:02,190
pieces of information that I just showed 
you. 

61
00:05:02,190 --> 00:05:06,570
So, one natural first attempt at this 
would be to define a smoothed model 

62
00:05:06,570 --> 00:05:11,607
that's very similar to the trigram models 
that we'd seen earlier in the class or 

63
00:05:11,607 --> 00:05:18,890
rather similar to the smoothing methods 
we'd seen for those models. 

64
00:05:18,890 --> 00:05:22,474
So I could take these nine different 
maximum likelihood estimates that I've 

65
00:05:22,474 --> 00:05:26,8
shown you here. 
So here, I have the regular trigram, 

66
00:05:26,8 --> 00:05:30,220
bigram, unigram estimate. 
Here, I condition on the word to that 

67
00:05:30,220 --> 00:05:33,920
being any, the condition on the last word 
being an adjective. 

68
00:05:33,920 --> 00:05:38,330
a condition on the author, the presence 
or absence in modeling the context. 

69
00:05:38,330 --> 00:05:41,9
The presence or absence of grammatical in 
the context. 

70
00:05:41,9 --> 00:05:44,669
I take all of these estimates and I 
simply take a linear interpolation of 

71
00:05:44,669 --> 00:05:49,485
these estimates, so I have now 
parameters, smoothing parameters. 

72
00:05:49,485 --> 00:05:55,373
This lambda 1, lambda 2, lambda 3, up to 
lambda 9, and these will all great link 

73
00:05:55,373 --> 00:05:59,240
to 0 and they sum to 1. 
Okay? 

74
00:05:59,240 --> 00:06:05,279
So, I wanted to go through this as a 
thought experiment, as a first way that 

75
00:06:05,279 --> 00:06:14,810
you might try to build a model that makes 
use of all this different information. 

76
00:06:14,810 --> 00:06:19,726
In practice, this kind of approach 
quickly becomes extremely unwieldy. 

77
00:06:19,726 --> 00:06:23,570
And, in fact, is beset by all kinds of 
practical problems, really makes it a 

78
00:06:23,570 --> 00:06:27,720
non-starter. 
So while this method, a trigram language 

79
00:06:27,720 --> 00:06:32,270
model that may use just these three 
estimates and these are three definitions 

80
00:06:32,270 --> 00:06:38,356
of contexts that works quite well. 
Once you try to extend these models to 

81
00:06:38,356 --> 00:06:43,700
include different types of feature, they 
become extremely unwieldy. 

82
00:06:43,700 --> 00:06:48,640
As we'll see, log linear models, a very 
clean and I think very elegant solution 

83
00:06:48,640 --> 00:06:53,428
to this problem of incorporating multiple 
sources of information in these 

84
00:06:53,428 --> 00:06:58,244
estimates. 
So here is a second, very important 

85
00:06:58,244 --> 00:07:03,726
motivating example for log linear models 
and that's the problem of tagging. 

86
00:07:03,726 --> 00:07:06,846
Specifically, in this example, we're 
going to look again at part of speech 

87
00:07:06,846 --> 00:07:10,781
tagging. 
So as a recap, the problem here is to 

88
00:07:10,781 --> 00:07:16,437
take some sentence, sentence, some 
sequence of words as input. 

89
00:07:16,437 --> 00:07:21,429
And to map this to a representation where 
each word has an associated tag, for 

90
00:07:21,429 --> 00:07:28,834
example, n for a noun, v for a verb, p 
for a preposition, and so on and so on. 

91
00:07:34,200 --> 00:07:39,572
So here is a natural estimation problem 
associated with this this particular 

92
00:07:39,572 --> 00:07:42,320
problem. 
Okay? 

93
00:07:42,320 --> 00:07:46,880
So if we take a particular word in a 
sentence, our task is going to be to 

94
00:07:46,880 --> 00:07:53,492
estimate a distribution over the possible 
tags at that position. 

95
00:07:53,492 --> 00:07:58,290
So again, that might be 40 or 50 possible 
part of speech tags. 

96
00:07:58,290 --> 00:08:02,715
So, ti is going to be the ith tag of the 
sequence, and we'll use wi to refer to 

97
00:08:02,715 --> 00:08:07,523
the ith word in the sequence. 
Okay? 

98
00:08:07,523 --> 00:08:11,303
So we're going to try to estimate the, 
the distribution over potential tags of 

99
00:08:11,303 --> 00:08:16,396
the ith position, that's t i. 
Now, under the definition of the problem 

100
00:08:16,396 --> 00:08:21,486
I'm going to give you here, we're 
going to condition on two things. 

101
00:08:21,486 --> 00:08:27,318
Firstly, we can potentially condition on 
any information in the entire sentence, 

102
00:08:27,318 --> 00:08:32,830
so w1 through wn. 
Here's the input sentence. 

103
00:08:32,830 --> 00:08:37,728
And secondly, we can condition on any 
information in the previous i minus 1 

104
00:08:37,728 --> 00:08:43,516
tags. 
So in this particular case, t1 through ti 

105
00:08:43,516 --> 00:08:54,0
minus 1 is equal to the tag sequence NNP 
up here or V, VB, DT, [UNKNOWN] JJ. 

106
00:08:54,0 --> 00:08:57,36
So that's actually a sequence of length 
five. 

107
00:08:57,36 --> 00:09:01,659
We're trying to estimate probability of 
t6, given this previous sequence of tags, 

108
00:09:01,659 --> 00:09:07,79
and the entire sentence as input. 
Okay, so this looks slightly different 

109
00:09:07,79 --> 00:09:12,563
from the kind of parameters you saw in 
hidden Markov models for tagging. 

110
00:09:12,563 --> 00:09:17,918
But we'll see a little later that if we 
can come up with accurate estimates of 

111
00:09:17,918 --> 00:09:23,188
these conditional probabilities, then 
they lead to a very direct and very 

112
00:09:23,188 --> 00:09:29,138
powerful tagging model, which is an 
alternative to the, the problem it well, 

113
00:09:29,138 --> 00:09:39,900
is an alternative to the hidden Markov 
models we've seen earlier in this course. 

114
00:09:39,900 --> 00:09:44,444
So again, the thing to realize here is 
that in coming up with this estimate, we 

115
00:09:44,444 --> 00:09:50,112
could look at all kinds of features of 
the history or context. 

116
00:09:50,112 --> 00:09:54,528
Okay, so if we look at this particular 
example here, where we're trying to tag 

117
00:09:54,528 --> 00:09:58,77
the word base. 
So we're trying to estimate a 

118
00:09:58,77 --> 00:10:01,80
distribution of the tags at this 
position. 

119
00:10:01,80 --> 00:10:04,536
We can look at various things to say 
we're trying to estimate the probability 

120
00:10:04,536 --> 00:10:07,992
of seeing the word base tagged as an NN, 
that's a common noun, singular noun in 

121
00:10:07,992 --> 00:10:12,852
English. 
We can condition on the fact that the 

122
00:10:12,852 --> 00:10:18,308
current word being tagged is the word 
base or we could condition on the fact 

123
00:10:18,308 --> 00:10:24,540
that the previous tag ti minus 1 is the 
tag JJ. 

124
00:10:24,540 --> 00:10:28,170
Or, we could condition on prefix or 
suffix information about the word being 

125
00:10:28,170 --> 00:10:31,807
tagged. 
So I could condition on the fact that wi 

126
00:10:31,807 --> 00:10:36,367
ends in the single letter e or the fact 
that wi ends in the pair of letters s 

127
00:10:36,367 --> 00:10:41,202
followed by e. 
I could condition on surrounding words, 

128
00:10:41,202 --> 00:10:43,966
in this case. 
I could condition on the fact the 

129
00:10:43,966 --> 00:10:48,318
previous word is important or I could 
condition on the fact that the next word 

130
00:10:48,318 --> 00:10:53,61
is the word from. 
So again, if we continue with this 

131
00:10:53,61 --> 00:10:57,733
example, anyone of these features of the 
previous context could be useful in 

132
00:10:57,733 --> 00:11:03,750
predicting the distribution over the tag 
at the ith position. 

133
00:11:03,750 --> 00:11:08,525
And we could, once again, come up with a 
method based on linear interpolation. 

134
00:11:08,525 --> 00:11:12,245
The old method for smoothing we saw 
within the context of trigram language 

135
00:11:12,245 --> 00:11:16,85
models, but it would quickly become 
really unwieldy as you incorporate more 

136
00:11:16,85 --> 00:11:19,667
and more sources of information. 

