1
00:00:00,012 --> 00:00:02,593
>> Okay.
So, that completes the lecture.

2
00:00:02,593 --> 00:00:06,641
So, just to summarize some main lessons
we've learned here.

3
00:00:06,641 --> 00:00:11,509
There were three steps in deriving the
language models I've shown you.

4
00:00:11,509 --> 00:00:16,927
The first step was to expand the joint
probability over a sequence of words, w1,

5
00:00:16,927 --> 00:00:20,550
w2, up to wn, using the Chain rule of
probabilities.

6
00:00:20,550 --> 00:00:25,850
This is what I showed you in the portion
of the lecture on Markov processes.

7
00:00:25,850 --> 00:00:32,217
And the second step was to make Markov
independence assumptions.

8
00:00:32,217 --> 00:00:38,751
In particular, assuming that the
probability of some word, wi, conditioned

9
00:00:38,751 --> 00:00:45,993
on the entire previous sequence of i minus
1 previous words, actually depends only on

10
00:00:45,993 --> 00:00:52,419
the previous two words in this sequence,
we call this a second-order Markov

11
00:00:52,419 --> 00:00:56,474
assumption.
And the final step in deriving this, these

12
00:00:56,474 --> 00:01:01,804
estimates was to smooth these diagram
estimates essentially using low order

13
00:01:01,804 --> 00:01:05,176
accounts.
And that was done either through the

14
00:01:05,176 --> 00:01:11,224
method of linear interpolation or through
this discounting method that I just showed

15
00:01:11,224 --> 00:01:14,387
you.
So, just briefly, language modelling is a

16
00:01:14,387 --> 00:01:20,067
huge industry and there's been a lot of
research in improved methods for language

17
00:01:20,067 --> 00:01:23,462
modelling.
Some areas of particular interest are

18
00:01:23,462 --> 00:01:29,150
methods that model the underlying topic of
documents or other long-range features of

19
00:01:29,150 --> 00:01:32,577
a document.
So, conditioning on just the previous two

20
00:01:32,577 --> 00:01:37,468
words is certainly limiting and in, in
some cases, you might want to condition on

21
00:01:37,468 --> 00:01:41,775
the fact that a, a document is about
sports, or is about politics, or the

22
00:01:41,775 --> 00:01:46,885
general topic that can influence the words
that are seen in the document and might be

23
00:01:46,885 --> 00:01:51,135
important to condition on that.
Or we might condition on words which are

24
00:01:51,135 --> 00:01:55,864
outside this two-word window, there's been
considerable interest in that problem.

25
00:01:55,864 --> 00:02:01,408
Another type of model we'll see later in
the class, is language models built based

26
00:02:01,408 --> 00:02:05,670
on syntactic models.
Language models that explicitly trying to

27
00:02:05,670 --> 00:02:10,598
incorporate grammatical information,
information about what sentences a

28
00:02:10,598 --> 00:02:14,112
grammatical versus non-grammatical in a
language.

29
00:02:14,112 --> 00:02:18,877
And again, these models can often capture
the long range features, which fall

30
00:02:18,877 --> 00:02:23,236
outside just a, a two-way window.
A rule though, it can be quite, quite

31
00:02:23,236 --> 00:02:28,809
difficult to improve upon language models.
They're simple, they're very efficient and

32
00:02:28,809 --> 00:02:31,440
they can get us a long way in many
problems.
