1
00:00:09,041 --> 00:00:10,386
Welcome back.

2
00:00:10,386 --> 00:00:14,627
In this module, we are going to
talk about text classification or

3
00:00:14,627 --> 00:00:17,440
supervised learning for text.

4
00:00:17,440 --> 00:00:23,030
To start and set up the case,
I want you to look at this paragraph.

5
00:00:23,030 --> 00:00:26,455
It is about a medical document.

6
00:00:26,455 --> 00:00:31,220
And I want you to think about
the specialty that this relates to.

7
00:00:31,220 --> 00:00:37,069
Is it nephrology, neurology, or podiatry?

8
00:00:37,069 --> 00:00:41,376
Looking at the words,
you would think that it's podiatry,

9
00:00:41,376 --> 00:00:43,488
because you have foot there,

10
00:00:43,488 --> 00:00:49,082
you have something to do with fungal
infection and skin infection, and so on.

11
00:00:49,082 --> 00:00:53,152
And, nephrology,
which is the science of kidneys or

12
00:00:53,152 --> 00:00:58,037
neurology that is about the brain,
This is more closely related

13
00:00:58,037 --> 00:01:02,032
to the study of foot, so
that's podiatry, right?

14
00:01:02,032 --> 00:01:06,513
Now, given the three classes that
you have already, that's nephrology,

15
00:01:06,513 --> 00:01:10,168
neurology and podiatry,
let's look at another paragraph.

16
00:01:10,168 --> 00:01:14,476
Here, you will see that it
talks about kidney failure.

17
00:01:14,476 --> 00:01:19,002
And just by looking at the first few
words you would know that this probably

18
00:01:19,002 --> 00:01:20,536
belongs to nephrology.

19
00:01:20,536 --> 00:01:23,856
Now think about how you
made that decision.

20
00:01:23,856 --> 00:01:27,840
Did you use the work nephrology anywhere?

21
00:01:27,840 --> 00:01:30,528
It's not there in text.

22
00:01:30,528 --> 00:01:32,880
Then how did you know
that it's nephrology?

23
00:01:34,480 --> 00:01:38,065
The icon is important, but when you
are talking about text classification,

24
00:01:38,065 --> 00:01:39,343
you don't have that icon.

25
00:01:39,343 --> 00:01:41,782
I just put it up there.

26
00:01:41,782 --> 00:01:47,506
So you're looking at the words, so kidney
failure or renal failure and you somehow

27
00:01:47,506 --> 00:01:53,090
magically know that nephrology relates
to kidneys and renal diseases and so on.

28
00:01:55,460 --> 00:02:02,028
This is an important characteristic
of identifying a class based on text.

29
00:02:02,028 --> 00:02:04,680
And what he just did is classification.

30
00:02:05,690 --> 00:02:07,818
So what is classification?

31
00:02:07,818 --> 00:02:10,159
You have a given set of classes,

32
00:02:10,159 --> 00:02:13,994
in this particular case we
had these three classes.

33
00:02:13,994 --> 00:02:20,908
And note that I don't even have to tell
you the name nephrology or neurology.

34
00:02:20,908 --> 00:02:24,990
I could just give them class one,
two, three or these three icons.

35
00:02:26,550 --> 00:02:28,929
But you know that they
are these three concepts,

36
00:02:28,929 --> 00:02:31,729
these three classes that you
want to classify them into.

37
00:02:31,729 --> 00:02:37,460
And then the task is to assign
the correct class label to a given input.

38
00:02:39,130 --> 00:02:44,433
There are a lot of examples of text
classification when you look at it.

39
00:02:44,433 --> 00:02:47,959
For example,
if you are looking at a news article, and

40
00:02:47,959 --> 00:02:51,564
depending on which page of
the newspaper it belongs to,

41
00:02:51,564 --> 00:02:56,190
you would want to categorize it as
politics, sports, or technology.

42
00:02:57,550 --> 00:02:59,790
There are many more classes, of course.

43
00:02:59,790 --> 00:03:00,857
But in this case,

44
00:03:00,857 --> 00:03:05,135
let's say we want to distinguish them
into one of these three classes.

45
00:03:05,135 --> 00:03:13,390
Or you get an email and you want to label
that email as a spam or not a spam.

46
00:03:13,390 --> 00:03:15,457
Should it go into spam folder?

47
00:03:15,457 --> 00:03:22,244
How does any mail client decide
that it should go in this category?

48
00:03:22,244 --> 00:03:27,999
Basically, what is happening is there is
a text classification model running time.

49
00:03:27,999 --> 00:03:29,970
Think of sentiment analysis.

50
00:03:29,970 --> 00:03:32,867
You read a movie review and
just by reading it,

51
00:03:32,867 --> 00:03:37,047
you want to decide if it's a positive
review or a negative review.

52
00:03:37,047 --> 00:03:41,030
The text classification can
actually be at very scales.

53
00:03:41,030 --> 00:03:45,306
All of these are really at the scale
of a document, and you could call

54
00:03:45,306 --> 00:03:50,199
a paragraph a document, or a news report
a document, or an email a document.

55
00:03:50,199 --> 00:03:55,270
But you could also have text
classification at a word level.

56
00:03:55,270 --> 00:03:57,927
So think of the problem
of spelling correction.

57
00:03:57,927 --> 00:04:01,446
And you have weather written two ways.

58
00:04:01,446 --> 00:04:06,650
One means the climate,
the other is a construct for the sentence.

59
00:04:08,110 --> 00:04:12,455
Which is the right spelling when
you are writing a sentence?

60
00:04:12,455 --> 00:04:14,100
The weather is great.

61
00:04:14,100 --> 00:04:18,133
Is it the first one or the second one?

62
00:04:18,133 --> 00:04:20,067
What about the word color and

63
00:04:20,067 --> 00:04:24,750
the way that it's spelled in
British English and American English?

64
00:04:26,800 --> 00:04:32,955
Depending on all the other words, and if
suppose you're writing a BBC news article,

65
00:04:32,955 --> 00:04:37,582
or reading one, you are most
likely to see the second spelling.

66
00:04:37,582 --> 00:04:42,642
The first spelling in that
context would appear as an error.

67
00:04:42,642 --> 00:04:45,170
You need to understand a correct for
those.

68
00:04:45,170 --> 00:04:47,429
That would be the spelling
correction problem.

69
00:04:47,429 --> 00:04:53,062
All of these tasks are in general what
are known as supervised learning tasks.

70
00:04:53,062 --> 00:04:58,906
It's supervised because just like
humans learn from past experiences,

71
00:04:58,906 --> 00:05:01,930
machines learn from past instances.

72
00:05:03,780 --> 00:05:07,589
So, for example,
in a supervised classification task,

73
00:05:07,589 --> 00:05:11,712
you have this training phase
where information is gathered and

74
00:05:11,712 --> 00:05:16,163
a model is built and an inference
phase where that model is applied.

75
00:05:16,163 --> 00:05:21,590
So, for example, you'll have a set of
inputs that we call the labeled input,

76
00:05:21,590 --> 00:05:25,619
where we know that this
particular instance is positive,

77
00:05:25,619 --> 00:05:28,060
this one is negative, and so on.

78
00:05:28,060 --> 00:05:30,629
So in this case, let's just do examples.

79
00:05:30,629 --> 00:05:37,403
Green and red, or light and
dark, or positive and negative.

80
00:05:37,403 --> 00:05:41,036
And you did that set that
the label set up instances and

81
00:05:41,036 --> 00:05:43,960
feed it into a classification algorithm.

82
00:05:45,240 --> 00:05:50,532
This classification algorithm
will learn which instances

83
00:05:50,532 --> 00:05:57,420
appear to be more positive than negative
and build a model for what it learns.

84
00:05:57,420 --> 00:06:01,969
Once you have the model,
you can used it in the inference phase,

85
00:06:01,969 --> 00:06:06,603
where you have unlabeled input and
then this model will take it and

86
00:06:06,603 --> 00:06:11,083
give out labels for those inputs,
for those input instances.

87
00:06:13,867 --> 00:06:20,460
In this supervised learning, you learn
a classification model on properties.

88
00:06:20,460 --> 00:06:25,474
So when we say instances,
you're really looking at properties,

89
00:06:25,474 --> 00:06:31,114
or features of these instances, and
the model that is learned is basically

90
00:06:31,114 --> 00:06:36,321
importance of sort, or weight,
that is given to those properties.

91
00:06:36,321 --> 00:06:39,882
And all of this is basically learned
from the labeled instances given to you.

92
00:06:42,307 --> 00:06:47,507
More formally, the set of attributes or
features that represents the input

93
00:06:47,507 --> 00:06:52,380
is denoted by x, usually written
as bold x, because it's a vector.

94
00:06:52,380 --> 00:06:57,577
It is a set of individual features,
and in this particular case

95
00:06:57,577 --> 00:07:02,978
let's say there are n features.What
would these n features be?

96
00:07:02,978 --> 00:07:07,875
Let's take an example of email and
in order to take with it spam or not,

97
00:07:07,875 --> 00:07:10,540
you would say where does it come from?

98
00:07:11,560 --> 00:07:16,860
Does it have Kind of interesting
words like Nigeria or

99
00:07:16,860 --> 00:07:21,705
prince or deposit money or
something like that?

100
00:07:21,705 --> 00:07:26,421
So those would be attributes on which
you make a decision whether this

101
00:07:26,421 --> 00:07:29,276
email should go in the spam folder or not.

102
00:07:29,276 --> 00:07:35,500
Then, you have the class label,
the set of class labels is Y.

103
00:07:36,970 --> 00:07:40,391
Let's say there are k classes.

104
00:07:40,391 --> 00:07:45,057
If it's just two classes like positive and
negative, then k will be 2.

105
00:07:45,057 --> 00:07:50,389
If it was this medical speciality
example we saw where it's nephrology,

106
00:07:50,389 --> 00:07:53,149
neurology or podiatry, then k is 3.

107
00:07:53,149 --> 00:07:58,446
One of those classes is
the class label assigned

108
00:07:58,446 --> 00:08:03,886
to the instance, and that is by small y.

109
00:08:03,886 --> 00:08:06,642
Once we have learned this
classification model,

110
00:08:06,642 --> 00:08:09,869
we apply that model to new
instances to predict the label.

111
00:08:12,370 --> 00:08:14,105
So when we look at these,

112
00:08:14,105 --> 00:08:19,242
there are some terminology of data
sets that you would see very commonly.

113
00:08:19,242 --> 00:08:23,650
So again, there are two phases, the
training phase and the inference phase.

114
00:08:23,650 --> 00:08:27,725
The training phase has labeled data set.

115
00:08:27,725 --> 00:08:32,313
And in general, in the inference
phase you have unlabeled data set.

116
00:08:32,313 --> 00:08:35,695
Unlabeled data set is where
you have all instances and

117
00:08:35,695 --> 00:08:39,070
you have the x defined,
but you don't have a y.

118
00:08:39,070 --> 00:08:40,472
You don't have a label.

119
00:08:40,472 --> 00:08:43,801
Whereas in the labeled data set,
you have the x and

120
00:08:43,801 --> 00:08:46,344
the y for every instance given to you.

121
00:08:48,886 --> 00:08:55,610
However, in training, you don't use the
entire label set for training purposes.

122
00:08:55,610 --> 00:08:59,094
Because then you will not
know how well your model is.

123
00:08:59,094 --> 00:09:03,538
So what you want to do is to use
a part of it as training data where

124
00:09:03,538 --> 00:09:05,940
you actually learn parameters.

125
00:09:05,940 --> 00:09:11,098
Learn the model, but
leave some aside as a validation data set,

126
00:09:11,098 --> 00:09:14,739
or it's sometimes called
hold out data set.

127
00:09:14,739 --> 00:09:20,156
So that in the training phase,
you can learn on the training data but

128
00:09:20,156 --> 00:09:25,680
then test, or evaluate, or
set parameters on the validation data.

129
00:09:27,060 --> 00:09:31,581
And then, you want another data set
to really test how well you do.

130
00:09:31,581 --> 00:09:34,457
We are to never use it in training.

131
00:09:34,457 --> 00:09:36,870
You don't set your
parameters based on that.

132
00:09:36,870 --> 00:09:41,355
But you just evaluate on that,
so that you can judge whether

133
00:09:41,355 --> 00:09:45,855
the model was really good or
not on completely unseen data.

134
00:09:45,855 --> 00:09:50,849
You have seen all of these concepts
in previous courses within this

135
00:09:50,849 --> 00:09:54,300
specialization that goes straight.

136
00:09:54,300 --> 00:09:55,920
But I want to kind of bring them here so

137
00:09:55,920 --> 00:09:58,581
that we have the context in
which we're going to talk about.

138
00:10:01,570 --> 00:10:04,999
One last thing about classification and
that's classification paradigms.

139
00:10:06,560 --> 00:10:13,139
We talked about cases when the set
of possible labels are two, right?

140
00:10:13,139 --> 00:10:19,546
Positive and negative or green and red or
yes and no or spam or not spam, right?

141
00:10:19,546 --> 00:10:23,896
All of these tasks are called
binary classification tasks because

142
00:10:23,896 --> 00:10:26,470
the number of possible classes is two.

143
00:10:27,620 --> 00:10:32,327
When that increases, When the number
of examples is more than two,

144
00:10:32,327 --> 00:10:35,712
number of classes I'm sorry,
is more than two,

145
00:10:35,712 --> 00:10:39,526
it's called multi-class
classification problem.

146
00:10:39,526 --> 00:10:47,052
And in some instances, you might want
to label with more than one labels.

147
00:10:47,052 --> 00:10:51,369
And when that happens, when the data and
classes are labeled by two or

148
00:10:51,369 --> 00:10:55,251
more labels,
that is called multi-label classification.

149
00:10:55,251 --> 00:10:58,608
Typically, we will look at
binary classification or

150
00:10:58,608 --> 00:11:03,907
multi-class classification, but there
are some instances within this module and

151
00:11:03,907 --> 00:11:08,175
the scores where we will look at
multi-label classification too.

152
00:11:08,175 --> 00:11:10,253
In general,
when we talk about classification,

153
00:11:10,253 --> 00:11:13,350
we're talking about binary or
multiclass classification.

154
00:11:13,350 --> 00:11:18,226
When you look at what questions to ask
in a supervised learning scenario, in

155
00:11:18,226 --> 00:11:23,731
the training phase, the questions that you
need to answer are what are the features?

156
00:11:23,731 --> 00:11:25,170
How do you represent them?

157
00:11:25,170 --> 00:11:28,925
How do you represent
the input given to you?

158
00:11:28,925 --> 00:11:33,727
What is a classification model or
the algorithm you're going to use?

159
00:11:33,727 --> 00:11:37,720
And what are the model parameters,
depending on the model you use?

160
00:11:39,670 --> 00:11:45,528
These are the questions that you will
answer while you're building your model.

161
00:11:45,528 --> 00:11:48,800
You need to know how do
you represent input?

162
00:11:48,800 --> 00:11:52,790
How are they going to train, or
what model are they going to train, and

163
00:11:52,790 --> 00:11:55,643
then what is the output
of that training process?

164
00:11:55,643 --> 00:11:57,566
And in the inference phase,

165
00:11:57,566 --> 00:12:01,901
you need to define what
are the expected performance measures?

166
00:12:01,901 --> 00:12:03,060
What is a good measure?

167
00:12:03,060 --> 00:12:06,312
How do you know that you
have built a great model?

168
00:12:06,312 --> 00:12:08,570
What if a performance measure
you'll use to determine that?

169
00:12:09,660 --> 00:12:13,888
And these are the questions you would
answer when you are building a supervised

170
00:12:13,888 --> 00:12:14,860
learning model.

171
00:12:14,860 --> 00:12:16,938
The same thing applies to text.

172
00:12:16,938 --> 00:12:20,610
And in the next few videos we
are going to answer these one by one.