1
00:00:00,000 --> 00:00:04,717
Hi, I'm Dan Jurafsky, and Chris Manning
and I are very happy to welcome you to our

2
00:00:04,717 --> 00:00:09,435
course on natural language processing.
This is a particularly exciting time to be

3
00:00:09,435 --> 00:00:13,979
working on natural language processing.
The vast amount of data on the Web and

4
00:00:13,979 --> 00:00:18,463
social media have made it possible to
build fantastic new applications. Let's

5
00:00:18,463 --> 00:00:23,007
look at one of them. Question answering.
You may know that IBM's Watson won the

6
00:00:23,007 --> 00:00:28,488
Jeopardy challenge on February sixteen,
2011. Answering questions like William

7
00:00:28,488 --> 00:00:34,974
Wilkinson's book inspired this author's
most famous novel. And you may know that

8
00:00:34,974 --> 00:00:42,239
the answer is Bram Stoker who famously
wrote [sound] Dracula. Another important

9
00:00:42,239 --> 00:00:46,515
task is information extraction. For
example, imagine that I have the following

10
00:00:46,515 --> 00:00:51,007
email from my colleague Chris about
scheduling a meeting. We'd like software

11
00:00:51,007 --> 00:00:56,385
to automatically notice that there are
dates, like tomorrow; times, like ten to

12
00:00:56,385 --> 00:01:01,343
eleven:30; in a room, like Gates 159;
extract those information, create a new

13
00:01:01,343 --> 00:01:06,442
calendar entry, and then populate a
calendar with this kind of structured

14
00:01:06,442 --> 00:01:11,541
information, with the event, date, start,
and end, for a calendar program. And

15
00:01:11,541 --> 00:01:18,996
modern email and calendar programs are
capable of doing this from text. Another

16
00:01:18,996 --> 00:01:23,508
application of this kind of information
extraction, involves sentiment analysis.

17
00:01:23,508 --> 00:01:28,134
Imagine that you're, interested in cameras
and you're reading a lot of reviews of

18
00:01:28,134 --> 00:01:32,034
cameras on the web, so here's a bunch of,
bunch of reviews. We'd like to

19
00:01:32,034 --> 00:01:35,999
automatically determine, from the reviews,
that what people care about in cameras,

20
00:01:35,999 --> 00:01:39,816
are particular attributes. If they're
buying a camera, they want to know if it

21
00:01:39,816 --> 00:01:43,731
has good zoom or affordability, or size
and weight. So, you want to automatically

22
00:01:43,731 --> 00:01:48,196
determine those attributes. And then we'd
like to automatically, for any particular

23
00:01:48,196 --> 00:01:52,629
attribute, determine how the reviewers
felt about those attributes. For example,

24
00:01:52,629 --> 00:01:57,006
if a reviewer said nice and compact to
carry, that's a positive sentiment, and

25
00:01:57,006 --> 00:02:01,383
here's another positive example. But a,
but a phrase like flimsy is a negative

26
00:02:01,383 --> 00:02:05,589
sentiment. So we'd like to automatically
detect for each sentence what the

27
00:02:05,589 --> 00:02:09,511
sentiment is, and then aggregate for each
feature for, say, presume for

28
00:02:09,511 --> 00:02:14,058
affordability, so we might decide that
this camera, the reviewers really like the

29
00:02:14,058 --> 00:02:19,257
flash. But they weren't so happy about the
ease of use. We might measure the positive

30
00:02:19,257 --> 00:02:23,730
and negative sentiment. About each
attribute and then aggregate those.

31
00:02:23,730 --> 00:02:28,485
Machine translation is another important
new application and machine translation

32
00:02:28,485 --> 00:02:33,240
can be fully automatic. So for example, we
might have a source sentence in Chinese

33
00:02:33,240 --> 00:02:37,995
and here's Stanford's phrasal MT system
translating that into English. But MT can

34
00:02:37,995 --> 00:02:42,926
also be used to help human translators. So
here we might have an Arabic text and the

35
00:02:42,926 --> 00:02:47,505
human translator translating it into
English might need some help from the MT

36
00:02:47,505 --> 00:02:52,084
system, for example, a collection of
possible next words that the MT system can

37
00:02:52,084 --> 00:02:57,254
build automatically and help the human
translator. Let's look at the state of the

38
00:02:57,254 --> 00:03:02,543
art in language technology. Like every
field, NLP's divided up into specialties

39
00:03:02,543 --> 00:03:07,591
and sub-specialties. A number of these
problems are pretty close to solved. So,

40
00:03:07,591 --> 00:03:12,901
for example, spam detection, while it's
very hard to completely detect spam in our

41
00:03:12,901 --> 00:03:17,752
email boxes, we don't have, 99 percent
spam, and that's because spam detection is

42
00:03:17,752 --> 00:03:22,926
a relatively, easy classification task. A
couple of important component tasks, part

43
00:03:22,926 --> 00:03:27,141
of speech tagging and named entity
tagging. We'll talk about those, later in

44
00:03:27,141 --> 00:03:30,978
the course. And those work at pretty high
accuracies. We're gonna get 97 percent

45
00:03:30,978 --> 00:03:34,977
accuracy in part of speech tagging, and
we'll see how that's important for

46
00:03:34,977 --> 00:03:39,025
parsing. In other tasks, we're making good
progress. Not as commercial not as

47
00:03:39,025 --> 00:03:43,276
completely solved but, there are systems
out there that are, that are being used.

48
00:03:43,276 --> 00:03:47,797
So we talked about sentiment analysis the
task of deciding, thumbs up or thumbs down

49
00:03:47,797 --> 00:03:51,595
on a sentence or a product. Component
technologies like word sense

50
00:03:51,595 --> 00:03:56,466
disambiguation deciding if we're talking
about a rodent or a computer mouse when

51
00:03:56,466 --> 00:04:01,397
people talk about mouses in a search.
We'll talk about parcing which is good

52
00:04:01,397 --> 00:04:05,606
enough now to be used in lots of
applications, and machine translation

53
00:04:05,606 --> 00:04:09,936
usable on the web. A number of
applications however are still quite hard.

54
00:04:09,936 --> 00:04:14,867
So for example, answering hard questions
like how effective is this medicine in

55
00:04:14,867 --> 00:04:19,798
treating that disease, by looking at the
web or by summarizing information we know

56
00:04:19,798 --> 00:04:24,557
is quite hard. Similarly, while we made
some progress on, deciding that, the

57
00:04:24,557 --> 00:04:30,067
sentence xyz company acquired abc company
yesterday means something similar to abc

58
00:04:30,067 --> 00:04:35,378
has been taken over by xyz. The general
problem of detecting that two phrases or

59
00:04:35,378 --> 00:04:40,260
sentences mean the same thing the
paraphrase tasks still quite hard. Even

60
00:04:40,260 --> 00:04:44,528
harder is the task of summarization,
reading a number of, let's say, news

61
00:04:44,528 --> 00:04:48,976
articles that say that the oh the
Dow Jones is up or the S&P500

62
00:04:48,976 --> 00:04:53,604
has jumped, and housing prices
rose, and aggregating that to give user

63
00:04:53,604 --> 00:04:58,336
information, like, in summary, the economy
is good. And finally, one of the hardest

64
00:04:58,336 --> 00:05:02,745
tasks in natural language processing:
carrying on a complete human-machine

65
00:05:02,745 --> 00:05:07,511
communication in dialogue. So, here's a
simple example asking about what movie is

66
00:05:07,511 --> 00:05:12,098
playing when and buy movie tickets, and
you can get applications that do that

67
00:05:12,098 --> 00:05:16,447
today. But the general problem of
understanding everything the user might

68
00:05:16,447 --> 00:05:22,162
ask for, and returning a sensible
response, is quite difficult. Why is

69
00:05:22,162 --> 00:05:27,328
natural language processing so difficult?
One cute example are the kinds of,

70
00:05:27,535 --> 00:05:33,115
ambiguity problems that are called crash
blossoms. So, ambiguity is any case where

71
00:05:33,115 --> 00:05:38,213
a surface form might have multiple
interpretations. A crash blossom is the

72
00:05:38,213 --> 00:05:43,724
name for a kind of headline that has two
meanings, and the ambiguity causes, a

73
00:05:43,724 --> 00:05:49,166
humorous interpretation. So, reading this
first headline, "Violinist Linked to JAL

74
00:05:49,166 --> 00:05:54,249
Crash Blossoms." You might think that the
main verb is linked and the violinist is

75
00:05:54,249 --> 00:05:59,342
being linked to what. He's being linked to
Japan Airline's crash blossoms. Well, what

76
00:05:59,342 --> 00:06:04,190
are crash blossoms? Well this headline
gave the name to this phenomenon because

77
00:06:04,190 --> 00:06:08,731
the actual interpretation that the
headline writer intended, the main verb

78
00:06:08,731 --> 00:06:13,456
was blossoms. Who does the blossoming, the
violinist, and this fact about being

79
00:06:13,456 --> 00:06:19,598
linked to JA crash was a modifier of
violinist. Similar kinds of syntactic

80
00:06:19,598 --> 00:06:25,045
ambiguities. So here "Teacher Strikes Idle
Kids", the writer intended the main verb

81
00:06:25,045 --> 00:06:30,288
to be idle. The strikes caused the kids to
be idle, but, of course, the humorous

82
00:06:30,288 --> 00:06:35,599
interpretation is that the teacher is
striking. Strike is the verb. And we have

83
00:06:35,599 --> 00:06:41,525
a teacher. Striking idle kids. Another
important kind of ambiguity, is word sense

84
00:06:41,525 --> 00:06:47,226
ambiguity. So in our third example, red
tape holds up new bridges, the writer

85
00:06:47,226 --> 00:06:53,232
intended holds up, to mean something like
delay. Call that sense one of holds up.

86
00:06:53,232 --> 00:06:59,389
But the amusing interpretation is the
second sense of holds up, which we might

87
00:06:59,389 --> 00:07:04,754
write down as to support. And now, we get
the interpretation that literal red-tape,

88
00:07:04,754 --> 00:07:09,641
as opposed to bureaucratic red-tape, is
actually supporting a bridge. And, we can

89
00:07:09,641 --> 00:07:14,900
see lots of other kinds of, ambiguities in
these actual headlines. Now, it turns out

90
00:07:14,900 --> 00:07:19,849
that it's not just amusing headlines that
have ambiguity. Ambiguity is pervasive

91
00:07:19,849 --> 00:07:24,798
throughout natural language text. Let's
look at a sensible, non-ambiguous-looking

92
00:07:24,798 --> 00:07:30,724
headline from the New York Times. So the
headline shortened it here a bit, is Fed

93
00:07:30,724 --> 00:07:37,088
raises interest rates, buy that seems
unambiguous. We have a verb here, a vital

94
00:07:37,088 --> 00:07:43,121
parser [inaudible] raises. What gets
raised? A noun phrase, a vital role to

95
00:07:43,121 --> 00:07:49,733
announce here interest rates. And we have
a verb phase, so raising interest rates

96
00:07:49,733 --> 00:07:54,877
and then we have the Fed. Make a little
noun phrase. And then we'll say, this is a

97
00:07:54,877 --> 00:07:59,073
sentence that has a noun phrase, Fed, and
a verb phrase, raises. And what gets

98
00:07:59,073 --> 00:08:03,214
raised is interest rates. So, this is
called a phrase structure parse. We'll

99
00:08:03,214 --> 00:08:09,110
talk about that, later in the course,
phrase structure. So, we could also write

100
00:08:09,110 --> 00:08:14,033
a dependency parse. So, we say the head
verb, raises, has an argument which is

101
00:08:14,033 --> 00:08:19,351
fed, and has another dependent, which is
rates. And, rates has another, itself has

102
00:08:19,351 --> 00:08:24,144
a dependent, interest. So, we can see the
main verb is raising. Well, another

103
00:08:24,144 --> 00:08:29,659
interpretation of the very same sentence,
one that people don't see but that parsers

104
00:08:29,659 --> 00:08:34,911
see right away, is that it's not raises
that's the main verb of the sentence, but

105
00:08:34,911 --> 00:08:40,229
interest. Somebody interests something,
and, that something that gets interested

106
00:08:40,229 --> 00:08:49,168
is rates. And what is interesting these
rates, well. It's fed raises, raises by

107
00:08:49,168 --> 00:08:53,507
the fed. So its a completely different
sentence with a different interpretation

108
00:08:53,507 --> 00:08:57,956
that something is interesting, the rates,
whatever that could mean, and it seems an

109
00:08:57,956 --> 00:09:02,460
unlikely interpretation for people. But of
course, for a parser, this is a perfectly

110
00:09:02,460 --> 00:09:06,634
reasonable interpretation that we have to
learn how to rule out. In fact, the

111
00:09:06,634 --> 00:09:11,193
sentence can get even more difficult. This
is, the actual headline was some, somewhat

112
00:09:11,193 --> 00:09:15,862
longer so we had fed raises interest rates
half a percent. Here we could imagine that

113
00:09:15,862 --> 00:09:20,421
rates is the verb and now we have what is
reading fed raises interest. The interest

114
00:09:20,421 --> 00:09:26,653
in federal raises. Are rating, half a
percent, so we might have a, a dependency

115
00:09:26,653 --> 00:09:31,988
structure like this. So again, interest.
Rates. The raises are what do the

116
00:09:31,988 --> 00:09:36,691
interesting and the Fed is a modifier of
raises. So, whether with our, phrase

117
00:09:36,691 --> 00:09:41,894
structure parse, or dependency parse, and
even more so as we add more words when get

118
00:09:41,894 --> 00:09:47,035
more and more ambiguity, that have to be
solved in order to build a parse, for each

119
00:09:47,035 --> 00:09:51,606
sentence. Now, the format of the course
you're going to have in video quizzes and

120
00:09:51,606 --> 00:09:55,464
most lectures will include a little quiz.
And they're there just to check basic

121
00:09:55,464 --> 00:09:59,273
understanding. They're simple multiple
choice questions. You can retake them if

122
00:09:59,273 --> 00:10:03,228
you get them wrong. Let's see one right
now. A number of other things make natural

123
00:10:03,228 --> 00:10:07,798
language understanding difficult. One of
them is the non standard English that we

124
00:10:07,798 --> 00:10:12,422
frequently see in, text like Twitter
feeds, where we have, capitalization and,

125
00:10:12,422 --> 00:10:17,230
unusual spelling of words, and hash tags
and user ID's and so on. So, all of our,

126
00:10:17,230 --> 00:10:22,285
parsers and part of speech taggers that
we're gonna make use of are often trained

127
00:10:22,285 --> 00:10:27,340
on very clean newspaper text English but,
the actual English in the, in the wild.

128
00:10:27,537 --> 00:10:32,680
Will cause us a lot of problems. We'll
have a lot of segmentation problems for

129
00:10:32,680 --> 00:10:38,151
example if we see that the string y o r k
dash any w as part as New York New Haven,

130
00:10:38,151 --> 00:10:43,573
how do we know, the correct segmentation
is New York? And New Haven. So the New

131
00:10:43,573 --> 00:10:49,678
York, New Haven railroad. And not
something like. York-dash-new. This word

132
00:10:49,678 --> 00:10:54,420
here is not a word like in-dash-law. We
have to solve the segmentation problem

133
00:10:54,420 --> 00:10:59,466
correctly. We have problems with idioms,
and with, new words that haven't be- seen

134
00:10:59,466 --> 00:11:04,147
before. And, we'll also have problems with
entity names, like the movie, A Bug's

135
00:11:04,147 --> 00:11:09,071
Life, which has English words in it, and
so it's often difficult to know where the

136
00:11:09,071 --> 00:11:13,995
movie name starts and ends. And this comes
up very often in biology. Where we have

137
00:11:13,995 --> 00:11:18,528
genes and proteins named with English
words. The task of natural understanding

138
00:11:18,528 --> 00:11:22,791
is very difficult. What tools do we need?
Well, we need knowledge about language,

139
00:11:22,791 --> 00:11:26,835
knowledge about the world and a way to
combine these knowledge sources. So

140
00:11:26,835 --> 00:11:31,207
generally the way we do this is to use
probabilistic models that are built from

141
00:11:31,207 --> 00:11:35,416
language data. So, for example, if we see
the word Maison, in French, we are very

142
00:11:35,416 --> 00:11:39,897
likely to translate that as the word house
in English. On the other hand if we see

143
00:11:39,897 --> 00:11:44,215
the word avocation all in French, we are
very unlikely to translate that as the

144
00:11:44,215 --> 00:11:48,539
general avocado. And training these
probabilistic models in general can be

145
00:11:48,539 --> 00:11:53,595
very hard. But it turns out that we can do
an approximate job of probabilistic models

146
00:11:53,595 --> 00:11:58,174
with rough text features and we'll
introduce those rough to, text features as

147
00:11:58,174 --> 00:12:02,358
we go on. So our goal in the class is
teaching key theory and methods for

148
00:12:02,358 --> 00:12:06,861
statistical natural language processing.
We'll talk about the Viterbi algorithm,

149
00:12:06,861 --> 00:12:11,307
nieve base, and maxen classifiers. We'll
introduce N gram language modeling and

150
00:12:11,307 --> 00:12:16,095
statistical parcing. We'll talk about the
inverted index and TFIDF and vector models

151
00:12:16,095 --> 00:12:20,256
of meaning that are important in
information retrieval. And we'll do this

152
00:12:20,256 --> 00:12:24,474
for practical, robust, real world
applications. We'll talk about information

153
00:12:24,474 --> 00:12:29,804
extraction, about spelling correction,
about information retrieval. The skills

154
00:12:29,804 --> 00:12:33,380
you'll need for the task, you'll need
simple linear algebra so you should know

155
00:12:33,380 --> 00:12:37,231
what a factor is and what a matrix is, you
should have some basic probability theory,

156
00:12:37,231 --> 00:12:40,990
and you need to know how to program an
either job over python because there'll be

157
00:12:40,990 --> 00:12:45,616
weekly programming assignments, you know
have your choice of languages. We're very

158
00:12:45,616 --> 00:12:49,561
happy to welcome you to our course on
Natural Language Processing and we look

159
00:12:49,561 --> 00:12:51,787
forward to seeing you in following
lectures.
