1
00:00:00,012 --> 00:00:05,624
Okay, so in this next segment of the class
we're going to talk about the IBM

2
00:00:05,624 --> 00:00:10,904
translation models.
The IBM models go back to the late 19,

3
00:00:10,904 --> 00:00:15,724
1980s, early 1990s.
They were seminal in really starting the

4
00:00:15,724 --> 00:00:19,890
whole statistical approach to the
translation problem.

5
00:00:19,890 --> 00:00:25,759
They form a central part of modern
statistical translation systems, and so

6
00:00:25,759 --> 00:00:31,067
we'll cover these models in some detail in
the segment of the class.

7
00:00:31,068 --> 00:00:37,737
So just to recap, this is a slide from the
last segment of this class.

8
00:00:37,738 --> 00:00:42,690
This is a recap of a noisy channel model.
Which was introduced by the IBM researches

9
00:00:42,690 --> 00:00:46,090
to translation.
So a noisy channel model, as we said, has

10
00:00:46,090 --> 00:00:49,288
two components.
Firstly, p of e, this is a language model,

11
00:00:49,288 --> 00:00:53,319
this is a model that assigns a probability
to any sentence in English.

12
00:00:53,320 --> 00:00:58,155
And secondly, p of f given e.
So this is what's called the translation

13
00:00:58,155 --> 00:01:01,027
model.
This assigns the model for a French

14
00:01:01,027 --> 00:01:04,530
sentence given an English sentence.
One note.

15
00:01:04,530 --> 00:01:09,882
Throughout this lecture, I'll follow
convention which is, I'll always assume

16
00:01:09,882 --> 00:01:15,089
that we're translating from a French
sentence to an English sentence.

17
00:01:15,090 --> 00:01:18,195
So that is our goal.
Okay.

18
00:01:18,195 --> 00:01:21,510
Of course, we might have other languages
for translating between.

19
00:01:21,510 --> 00:01:25,930
But for the purposes of this lecture,
we'll always assume we're translating from

20
00:01:25,930 --> 00:01:28,902
French to English.
Okay, so we have a language model for

21
00:01:28,902 --> 00:01:32,014
strings in English.
We have a translation model which, as I

22
00:01:32,014 --> 00:01:35,413
said, is in a sense backwards.
It says what's the conditional probability

23
00:01:35,413 --> 00:01:37,582
of any French sentence, given an English
sentence.

24
00:01:37,582 --> 00:01:43,939
When we translate on to this model we
search for the English sentence that

25
00:01:43,939 --> 00:01:48,007
maximizes the product of p of e and p of f
given e.

26
00:01:48,008 --> 00:01:51,679
And so when we evaluate for particular
sentence f.

27
00:01:51,680 --> 00:01:56,608
Different English translations, we take
both of these terms into account.

28
00:01:56,608 --> 00:02:03,038
So, the roadmap for the next few lectures
in this course on translation is as

29
00:02:03,038 --> 00:02:06,784
follows.
I'm first going to describe IBM Models 1

30
00:02:06,784 --> 00:02:09,968
and 2.
So IBM went through a series of models, I

31
00:02:09,968 --> 00:02:15,052
think 1 through 5, and we're just going to
do the first two because they will

32
00:02:15,052 --> 00:02:18,771
essentially introduce many of the
important ideas.

33
00:02:18,771 --> 00:02:23,855
And actually model 2 is a pretty decent
model for the problem we're going to be

34
00:02:23,855 --> 00:02:27,435
looking at.
We'll then go on to describe phrase-based

35
00:02:27,435 --> 00:02:30,840
models.
So phrased-based models were invented

36
00:02:30,840 --> 00:02:36,688
around the late 1990s, and they're in some
sense a second generation of s-, school

37
00:02:36,688 --> 00:02:40,537
translation systems.
This is the first generation.

38
00:02:40,538 --> 00:02:45,441
They make direct use of ideas from these
IBM models, and so these IBM models will,

39
00:02:45,441 --> 00:02:49,675
in some sense, form a basis of
phrase-based models, but phrase-based

40
00:02:49,675 --> 00:02:52,549
models work much, much better than IBM
models.

41
00:02:52,550 --> 00:02:56,736
They do, in fact form the basis of many
modern statistical translation systems, so

42
00:02:56,736 --> 00:03:00,120
for example, Google translate is heavily
built on this technology.

43
00:03:00,120 --> 00:03:09,100
So, we'll first talk about IBM Model 1 and
then we'll talk about IBM Model 2 and then

44
00:03:09,100 --> 00:03:16,673
finally we'll talk about parameter
estimation in models 1 and 2.

45
00:03:16,674 --> 00:03:18,964
How do we actually learn the parameters of
these models?

46
00:03:18,965 --> 00:03:24,053
From data consisting of example
translations.
