1
00:00:09,555 --> 00:00:14,352
In this video, we are going go in
a little bit more detail about how

2
00:00:14,352 --> 00:00:19,063
some of what we have learned in
this course can be all put together

3
00:00:19,063 --> 00:00:23,100
in this specific natural
language processing tasks.

4
00:00:23,100 --> 00:00:29,499
That are really important in a lot
of tasks of understanding free-text.

5
00:00:29,499 --> 00:00:32,870
Information is,
as we have seen so many times,

6
00:00:32,870 --> 00:00:37,070
is hidden in free-text in
very interesting ways.

7
00:00:37,070 --> 00:00:41,460
You have most traditional traditional
information that is structured,

8
00:00:41,460 --> 00:00:45,840
while a lot of information now is
in unstructured free text form.

9
00:00:48,900 --> 00:00:52,250
As we have seen earlier when I was
introducing this whole course,

10
00:00:52,250 --> 00:00:56,720
that about 80% of data is
now in unstructured form,

11
00:00:56,720 --> 00:01:00,320
in blogs, in websites,
on websites, and so on.

12
00:01:01,510 --> 00:01:02,580
And it's growing.

13
00:01:02,580 --> 00:01:03,707
You have tweets and

14
00:01:03,707 --> 00:01:07,590
other social media that is kind
of adding more information here.

15
00:01:09,390 --> 00:01:10,990
So the question then becomes,

16
00:01:10,990 --> 00:01:15,370
how do you convert this unstructured
text to structured form?

17
00:01:15,370 --> 00:01:19,060
You don't necessarily have to convert it
all, but more of what im trying to get

18
00:01:19,060 --> 00:01:25,680
here is, how do you extract relevent
information from unstructured text.

19
00:01:25,680 --> 00:01:29,350
And if you want to make it searchable or
make it usable later on you would probably

20
00:01:29,350 --> 00:01:32,960
put it in a structured form as
it has been traditionally kept.

21
00:01:32,960 --> 00:01:36,880
So that way it is conversion from
unstructured to structured form text.

22
00:01:38,440 --> 00:01:40,330
And that's where information
extraction comes in.

23
00:01:41,470 --> 00:01:46,520
The goal of this task is to identify and
extract fields of interest from free text.

24
00:01:47,880 --> 00:01:52,450
So let's take an example
of a medical article.

25
00:01:52,450 --> 00:01:55,860
This one is from WebMD
from quite some time back.

26
00:01:55,860 --> 00:02:01,189
But it says Erbitux helps
treat advanced lung cancer.

27
00:02:01,189 --> 00:02:06,096
By looking at this, you know that one
of the pieces of information that

28
00:02:06,096 --> 00:02:09,102
you want to extract would be the headline.

29
00:02:09,102 --> 00:02:13,925
But very importantly also,
the components of that saying that

30
00:02:13,925 --> 00:02:17,930
Erbitux is a particular treatment for
lung cancer.

31
00:02:19,730 --> 00:02:25,012
And the fact that advance lung cancer
is a specific form of lung cancer.

32
00:02:25,012 --> 00:02:29,550
So you're going to understand
which modifier could be dropped.

33
00:02:29,550 --> 00:02:33,100
So in this case, advanced could
be taught but lung couldn't be.

34
00:02:33,100 --> 00:02:37,920
Because lung cancer treatments are
different from let's say blood cancer or

35
00:02:38,980 --> 00:02:40,060
breast cancer and so forth.

36
00:02:42,230 --> 00:02:46,490
But this webpage has more
information than just the headline.

37
00:02:47,650 --> 00:02:54,220
It has some information about who wrote
this article or who reviewed this article.

38
00:02:55,670 --> 00:03:02,310
Or when was it published or
the place it was published from and so on.

39
00:03:02,310 --> 00:03:07,460
And you can see that this happens all
the time in news reports and other places

40
00:03:07,460 --> 00:03:13,520
when you see some piece of information,
some news article or a tweet.

41
00:03:13,520 --> 00:03:18,440
It has more than just the text
that is important, and

42
00:03:18,440 --> 00:03:20,470
are of some interest for people.

43
00:03:20,470 --> 00:03:23,730
So who wrote a particular news report?

44
00:03:23,730 --> 00:03:24,710
Where is it coming from?

45
00:03:26,320 --> 00:03:27,550
What does it say?

46
00:03:27,550 --> 00:03:31,062
Who or what are the other
persons involved in this story?

47
00:03:31,062 --> 00:03:37,148
This is the typical representation
of an information extraction task.

48
00:03:39,516 --> 00:03:43,230
The first thing to notice here is
something about fields of interest.

49
00:03:44,680 --> 00:03:47,309
These fields of interest
are named entities.

50
00:03:49,350 --> 00:03:54,504
If it's news, we are talking about people,
and places, and date, and

51
00:03:54,504 --> 00:03:59,220
organizations, and
other geopolitical entities.

52
00:03:59,220 --> 00:04:00,560
When you say White House,

53
00:04:00,560 --> 00:04:04,710
you mean something completely
different than a white house.

54
00:04:04,710 --> 00:04:09,538
Or when you're saying 10 Downing Street
it's a completely different

55
00:04:09,538 --> 00:04:13,189
meaning than the actual
address 10 Downing Street.

56
00:04:13,189 --> 00:04:18,368
If you're talking about finance you're
talking about money, or monetary values.

57
00:04:18,368 --> 00:04:23,987
About how much worth particular save for,
or the names of companies,

58
00:04:23,987 --> 00:04:28,000
or the stock price index
of a particular company.

59
00:04:29,820 --> 00:04:32,780
When talking about medicine and health,
you're talking about diseases and

60
00:04:32,780 --> 00:04:34,200
drugs and procedures.

61
00:04:35,230 --> 00:04:40,030
And especially if you're talking about the
protected information, protected health

62
00:04:40,030 --> 00:04:46,260
information, then you're talking about
address and their unique identifiers.

63
00:04:46,260 --> 00:04:51,020
Sometimes even emails and
URLs and provisions and so on.

64
00:04:51,020 --> 00:04:55,039
So you can see that
these name can be very,

65
00:04:55,039 --> 00:05:01,178
very diverse as they don't all follow
the same ideas and the same features.

66
00:05:01,178 --> 00:05:05,577
So for example,
in news it's fairly common for people,

67
00:05:05,577 --> 00:05:08,338
name and places to be capitalized.

68
00:05:08,338 --> 00:05:13,402
So you have a title case there,
you would see China

69
00:05:13,402 --> 00:05:17,871
with C capital more
often than with small c.

70
00:05:17,871 --> 00:05:20,829
You have dates,
they are typically of a particular format,

71
00:05:20,829 --> 00:05:23,680
we have seen some examples of
that early on in this course.

72
00:05:24,790 --> 00:05:27,830
But those don't apply in medicine.

73
00:05:27,830 --> 00:05:31,850
So, for example, when you're talking about
lung cancer, you don't have capital L and

74
00:05:31,850 --> 00:05:32,520
capital C.

75
00:05:34,410 --> 00:05:39,410
When you're writing it out,
it's going to be all small case.

76
00:05:39,410 --> 00:05:44,620
So the particular rule that you might have
for named entity being title cased is not

77
00:05:44,620 --> 00:05:47,229
applicable for some of the data sets,
like in medicine.

78
00:05:49,650 --> 00:05:54,040
Once you have understood what named
entities are, the next is relations.

79
00:05:54,040 --> 00:05:59,790
Relations are, basically is, what
happened to who, when, where, and so on.

80
00:05:59,790 --> 00:06:05,700
So again, here even though the particular
pieces of information are named entities,

81
00:06:05,700 --> 00:06:10,100
so who is basically a person or
an organization.

82
00:06:10,100 --> 00:06:12,640
When is a particular time, or a date.

83
00:06:12,640 --> 00:06:18,790
Where is a place, but it also has this
relationship between these named entities,

84
00:06:18,790 --> 00:06:21,250
to get meaningful semantics out of it.

85
00:06:22,700 --> 00:06:26,880
So let's take these two examples in
more detail and see what that means.

86
00:06:26,880 --> 00:06:31,940
So named entity recognition relies
on something called named entities.

87
00:06:31,940 --> 00:06:35,230
Named entities are noun phrases
that are of specific type and

88
00:06:35,230 --> 00:06:38,650
refer to specific individuals,
places, organizations, and so on.

89
00:06:39,740 --> 00:06:45,850
And the named entity recognition task
is a set of techniques and methods

90
00:06:45,850 --> 00:06:52,170
that would help identify all mentions
of predefined named entities in text.

91
00:06:52,170 --> 00:06:56,830
So for example,
you are to identify the mention,

92
00:06:56,830 --> 00:06:59,680
the actual phrase where
something is happening.

93
00:06:59,680 --> 00:07:04,280
So you need to know where does that
mention start and where does it end?

94
00:07:04,280 --> 00:07:09,480
So this is the boundary detection
subtask within named entity recognition.

95
00:07:09,480 --> 00:07:14,970
Once you understood the boundary,
the next task is to type it,

96
00:07:14,970 --> 00:07:19,730
to tag it, or to classify one of the named
entity types that you are interested in.

97
00:07:21,170 --> 00:07:26,593
So for example,
if you have the word Chicago,

98
00:07:26,593 --> 00:07:34,060
it could be a place,
it could be an album, it could be a font.

99
00:07:34,060 --> 00:07:37,190
So depending on what
kind of variation it is,

100
00:07:37,190 --> 00:07:41,670
you need to know what is the label
it should be assigned to this word.

101
00:07:41,670 --> 00:07:46,750
Once you identify that Chicago
is your element phrase.

102
00:07:46,750 --> 00:07:49,710
So let's take an example
of an named entity task

103
00:07:49,710 --> 00:07:52,100
in the context of a medical document.

104
00:07:53,850 --> 00:07:56,810
I'm going to leave this one and
let you try out

105
00:07:56,810 --> 00:08:00,510
what are the different named entities
here that you would be interested in.

106
00:08:03,300 --> 00:08:07,560
Okay now that you've had a chance
to see what are the named entities,

107
00:08:07,560 --> 00:08:08,770
let's see some of them.

108
00:08:11,030 --> 00:08:16,168
One big one is the disease or
the condition or the symptom, right?

109
00:08:16,168 --> 00:08:19,705
So you have bilateral hand numbness or

110
00:08:19,705 --> 00:08:24,467
occasional weakness or
C5-6 disc herniation.

111
00:08:24,467 --> 00:08:27,800
Okay, but you also have age here.

112
00:08:27,800 --> 00:08:32,690
So you could say age is 63 or
you could say age is 63 year old.

113
00:08:33,850 --> 00:08:36,880
You have gender, female.

114
00:08:36,880 --> 00:08:44,441
You have mention of a medical
specialty in neurologist.

115
00:08:44,441 --> 00:08:49,874
You have a procedure that was done,
so you have a workup or an MRI.

116
00:08:49,874 --> 00:08:55,882
And in relation to all of this, you
also have body parts like feet and hand.

117
00:08:55,882 --> 00:09:00,654
And certainly, you come to a conclusion
where you can see that these

118
00:09:00,654 --> 00:09:04,910
named entities don't have to
be separate from each other.

119
00:09:04,910 --> 00:09:13,380
So for example, bilateral hand numbness is
a completely valid named entity in itself.

120
00:09:13,380 --> 00:09:16,860
But hand is also a named
entity that is of interest.

121
00:09:16,860 --> 00:09:18,190
So then the question becomes,

122
00:09:18,190 --> 00:09:22,390
what is the granularity with
which you need to annotate this.

123
00:09:24,160 --> 00:09:26,460
And if you're doing hand
as a separate annotation,

124
00:09:26,460 --> 00:09:32,350
why not C5-6 disc or cord because you're
talking about spinal cord in this case,

125
00:09:32,350 --> 00:09:33,710
you're talking about an anatomical part.

126
00:09:34,970 --> 00:09:38,680
So these are the decisions that need to
be made when you're defining this named

127
00:09:38,680 --> 00:09:40,050
entity task.

128
00:09:40,050 --> 00:09:43,970
So in this particular case,
if you define the task to be one,

129
00:09:43,970 --> 00:09:49,880
where you say we are only interested
in the condition, the diagnosis,

130
00:09:49,880 --> 00:09:56,190
the age, the gender, the procedures
that were done, and the specialty.

131
00:09:56,190 --> 00:09:57,070
That's it.

132
00:09:57,070 --> 00:10:02,050
In this then, you're not going to focus
on a body part, hand and feet, and

133
00:10:02,050 --> 00:10:03,641
even if they are there, that's fine.

134
00:10:03,641 --> 00:10:06,320
And some decision already
was done that way.

135
00:10:06,320 --> 00:10:12,442
So for example, nobody said that patient
is not a valid entity in this role,

136
00:10:12,442 --> 00:10:15,324
it could be, it very well could be.

137
00:10:15,324 --> 00:10:19,131
If neurologist is a valid entity,
patient could be a valid entity, and

138
00:10:19,131 --> 00:10:23,650
that's a decision that has to be made when
you're defining this named entity task.

139
00:10:24,720 --> 00:10:29,010
Suppose you have these goals then,
you know that you have not only

140
00:10:29,010 --> 00:10:32,780
identified the boundary, but
also colored it in some way,

141
00:10:32,780 --> 00:10:38,020
tagged it to see what is the type of this
[INAUDIBLE] that we have identified.

142
00:10:39,850 --> 00:10:43,160
Okay, now that we know the task well,

143
00:10:43,160 --> 00:10:46,500
what are the approaches that people
use to identify named entities?

144
00:10:47,600 --> 00:10:50,950
It really depends on the kind
of entities that we have.

145
00:10:50,950 --> 00:10:55,090
So for example, if you have
well-formatted fields like dates and

146
00:10:55,090 --> 00:10:58,930
phone numbers,
then you would use regular expressions.

147
00:10:58,930 --> 00:11:02,650
Recall that we have done the regular
expressions for date, for

148
00:11:02,650 --> 00:11:05,440
example, in week one of this course.

149
00:11:05,440 --> 00:11:10,010
And that is a very valid named
entity identification task.

150
00:11:10,010 --> 00:11:13,500
In fact,
the assignment was really asking you to

151
00:11:13,500 --> 00:11:18,330
do an information extraction task for
dates from the given text file.

152
00:11:20,760 --> 00:11:24,870
For other fields, it's fairly common
to use a machine learning approach.

153
00:11:24,870 --> 00:11:29,040
In fact, even for dates and
phone numbers you might want to

154
00:11:29,040 --> 00:11:33,430
use a machine learning approach, where you
use these regular expressions as features.

155
00:11:33,430 --> 00:11:37,930
But you have other features that
help them define this particular

156
00:11:39,140 --> 00:11:40,720
extraction to be valid.

157
00:11:40,720 --> 00:11:44,850
So for example, phone numbers and
fax numbers are very similar.

158
00:11:46,230 --> 00:11:50,720
The only thing that differentiates them
is what is the number preceded by.

159
00:11:50,720 --> 00:11:54,593
So is it phone dot, or is it fax dot.

160
00:11:54,593 --> 00:11:58,165
So some partition information
kind of helps doing that, and

161
00:11:58,165 --> 00:12:02,717
you could either include it as a regular
expression or let those be features in

162
00:12:02,717 --> 00:12:06,656
a machine learning model that you
then train over large data sets.

163
00:12:06,656 --> 00:12:16,250
The standard NER task in natural language
processing is this four-class model.

164
00:12:16,250 --> 00:12:23,020
Person, organization, location, and
everything else, or everything outside.

165
00:12:23,020 --> 00:12:28,790
So, if you have a sentence like,
John met Brenda.

166
00:12:29,790 --> 00:12:31,660
You want to say that John and

167
00:12:31,660 --> 00:12:37,540
Brenda are named entities so
you have those two persons.

168
00:12:37,540 --> 00:12:41,830
And then Matt is a word that is
outside of any named entity, so

169
00:12:41,830 --> 00:12:43,928
you are going to use these two classes.

170
00:12:43,928 --> 00:12:48,460
So you're going to say John is a person,
Matt is outside and

171
00:12:48,460 --> 00:12:49,669
Brenda is a person again.

172
00:12:51,910 --> 00:12:54,550
Once you have identified these,
then you are going to go and

173
00:12:55,580 --> 00:12:58,090
ask the question about relation and
relation extraction.

174
00:12:59,200 --> 00:13:02,142
The relation extraction task is
identifying the relationship between

175
00:13:02,142 --> 00:13:02,737
named entities.

176
00:13:04,200 --> 00:13:06,640
So let's take an example of

177
00:13:06,640 --> 00:13:10,840
something that we have seen earlier
today on Erbitux help treat lung cancer.

178
00:13:12,270 --> 00:13:18,170
This sentence has a relationship
between two named entities.

179
00:13:18,170 --> 00:13:19,880
One is Erbitux.

180
00:13:19,880 --> 00:13:24,460
Erbitux is colored yellow here to
represent that it is a diagnosis,

181
00:13:24,460 --> 00:13:27,330
sorry it is a treatment of some kind.

182
00:13:27,330 --> 00:13:30,830
And then you have lung
cancer that is represented,

183
00:13:30,830 --> 00:13:34,300
that's in green to represent
a different name identity.

184
00:13:34,300 --> 00:13:37,430
In this particular case,
a disease or a diagnosis.

185
00:13:39,500 --> 00:13:41,860
And then you have
a relationship between them.

186
00:13:41,860 --> 00:13:45,709
So Erbitux and lung cancer
are linked by a treatment relation.

187
00:13:45,709 --> 00:13:50,675
Going from Erbitux to lung cancer,
that says Erbitux is a treatment for

188
00:13:50,675 --> 00:13:51,677
lung cancer.

189
00:13:51,677 --> 00:13:56,750
You could easily have
a link going the other way.

190
00:13:56,750 --> 00:14:02,880
Lung cancer and to Erbitux,
lung cancer is treated by Erbitux.

191
00:14:02,880 --> 00:14:06,000
So this is the relation,
a very simple relation,

192
00:14:06,000 --> 00:14:11,330
a binder relation between a disease and
a treatment, or a drug.

193
00:14:14,370 --> 00:14:17,320
The other task is co-reference resolution.

194
00:14:18,580 --> 00:14:21,880
That is to disambiguate
mentions in text and

195
00:14:21,880 --> 00:14:25,500
group mentions together if they
are referring to the same entity.

196
00:14:26,760 --> 00:14:30,590
An example would be if Anita
met Joseph at the market and

197
00:14:30,590 --> 00:14:37,270
he surprised her with a rose, then you
have two named entities, Anita and Joseph.

198
00:14:37,270 --> 00:14:40,270
But then,
the second sentence uses pronouns,

199
00:14:40,270 --> 00:14:44,110
he to refer to Joseph and
her to refer to Anita.

200
00:14:45,175 --> 00:14:51,510
So in this case, it's pronoun resolution
where you are making an inference

201
00:14:51,510 --> 00:14:56,910
that if Anita met Joseph at the market,
Joseph surprised Anita with a rose.

202
00:14:59,360 --> 00:15:00,950
Why is that important?

203
00:15:00,950 --> 00:15:06,650
That is in a question answering task
where you're given a question and

204
00:15:06,650 --> 00:15:11,650
you need to find the most
appropriate answer from text.

205
00:15:11,650 --> 00:15:15,890
Now appropriateness is something that
we can talk about in more detail.

206
00:15:15,890 --> 00:15:19,330
But let's say in this case,
the correct answer, or

207
00:15:19,330 --> 00:15:22,600
the set of correct answers,
would be one definition here.

208
00:15:24,715 --> 00:15:29,290
Andyou would need,
both named entity recognition and

209
00:15:29,290 --> 00:15:31,770
relation extraction to
answer these questions.

210
00:15:31,770 --> 00:15:35,470
So, for example,
the question was Erbitux treat?

211
00:15:35,470 --> 00:15:39,550
You first have to identify
that Erbitux is a treatment.

212
00:15:39,550 --> 00:15:41,392
The relation is treat and

213
00:15:41,392 --> 00:15:46,741
then somehow fill a slot in this relation
to say Erbitux is a treatment for

214
00:15:46,741 --> 00:15:51,145
something and that comes from text and
that is lung cancer.

215
00:15:51,145 --> 00:15:53,111
Or who gave Anita the rose,

216
00:15:53,111 --> 00:15:57,986
where you are doing some sort of
pronoun resolution to then identify

217
00:15:57,986 --> 00:16:03,129
that Joseph was who was the person who
gave Anita the rose at the market.

218
00:16:03,129 --> 00:16:05,844
So this builds on the named
entity recognition task and

219
00:16:05,844 --> 00:16:09,990
the relation extraction and co-reference
resolution that we have seen earlier.

220
00:16:12,890 --> 00:16:18,800
So, to conclude, we see that Information
Extraction is important task for

221
00:16:18,800 --> 00:16:22,910
natural language understanding and
making sense of textual data.

222
00:16:22,910 --> 00:16:26,970
It is the first step in
converting this unstructured

223
00:16:26,970 --> 00:16:28,650
text in to more structured form.

224
00:16:30,010 --> 00:16:34,240
Name Entity Recognition becomes a key
building block in addressing these

225
00:16:34,240 --> 00:16:38,535
tasks and these advanced NLP tasks.

226
00:16:38,535 --> 00:16:43,955
And named Entity Recognition systems
use the supervised machine learning

227
00:16:43,955 --> 00:16:48,395
approaches and text mining approaches
that we have discussed in the course.

228
00:16:48,395 --> 00:16:52,655
So for example, if the entity that
you need to recognize is a date,

229
00:16:52,655 --> 00:16:58,710
you are using typically expressions
that we've talked about in week one.

230
00:16:58,710 --> 00:17:03,710
If you are talking about
extracting people name.

231
00:17:03,710 --> 00:17:08,220
So person name or organization name,
you are not only using emotional learning

232
00:17:08,220 --> 00:17:12,640
model to identify what is an identity and

233
00:17:12,640 --> 00:17:17,360
what label you should get it, but
also the features that you're going to use

234
00:17:17,360 --> 00:17:20,910
are coming from what we
talked about in week two.

235
00:17:20,910 --> 00:17:25,240
So, for example, we want to know that,
yes, if it is capitalized or

236
00:17:25,240 --> 00:17:28,840
not, but what is the part of
speech of a particular word?

237
00:17:28,840 --> 00:17:30,271
Is it a noun or a verb?

238
00:17:30,271 --> 00:17:34,221
What is the semantic
role that a particular

239
00:17:34,221 --> 00:17:38,730
word is playing in a given
context in a sentence?

240
00:17:38,730 --> 00:17:42,650
And these could be features that you
would then put in in a named entity

241
00:17:42,650 --> 00:17:43,873
recognition model.

242
00:17:43,873 --> 00:17:49,234
NLTK has an in-built NER
model that does trained,

243
00:17:49,234 --> 00:17:54,595
or new datasets and so on,
for the standard task for

244
00:17:54,595 --> 00:18:00,320
the person, organization, location task.

245
00:18:00,320 --> 00:18:05,320
But I have shown in this video
the named entity recognition

246
00:18:05,320 --> 00:18:07,990
problem goes beyond just news.

247
00:18:07,990 --> 00:18:10,450
If your even extending it to finance or

248
00:18:10,450 --> 00:18:15,760
then if you go all the way to medicine,
it's a completely open problem.

249
00:18:15,760 --> 00:18:18,600
Where the topics that we
have talked about in this

250
00:18:18,600 --> 00:18:22,710
course kind of puts it all together and
brings it all together.

251
00:18:22,710 --> 00:18:26,190
Hope you had fun with the course and
I had fun presenting it to you.