1
00:00:01,660 --> 00:00:02,250
Hi.

2
00:00:02,250 --> 00:00:03,720
My name is Ciro Donalek and

3
00:00:03,720 --> 00:00:07,220
this is the first of the three
modules about classification.

4
00:00:08,990 --> 00:00:13,190
Today, I recap the basic concepts
in supervised learning and

5
00:00:13,190 --> 00:00:18,830
then I'll show how to measure the accuracy
of the classifier and how to evaluate it.

6
00:00:18,830 --> 00:00:22,180
In the next classes we'll see some model,
some modern details.

7
00:00:24,640 --> 00:00:27,390
Classification is where we're
[INAUDIBLE] it's one of

8
00:00:27,390 --> 00:00:30,230
the most popular data mining tasks.

9
00:00:30,230 --> 00:00:34,820
In a classification problem, we want
to divide our samples in classes and

10
00:00:34,820 --> 00:00:38,090
assign each object to a given class.

11
00:00:38,090 --> 00:00:42,520
Because the class label of each
training example is provided,

12
00:00:42,520 --> 00:00:45,480
the learning process is
said to be supervised.

13
00:00:46,660 --> 00:00:51,780
And as we have seen in the introduction to
machine learning, we talk about supervised

14
00:00:51,780 --> 00:00:56,090
learning when, for some examples,
the correct results are known and

15
00:00:56,090 --> 00:01:00,200
are given in input to the model
during the learning process.

16
00:01:00,200 --> 00:01:05,640
Ideally the classifier will be then
able to generalize, that means

17
00:01:05,640 --> 00:01:11,150
it will be able to determine the correct
results also for unseen distances.

18
00:01:14,070 --> 00:01:19,360
We have also talked about the data sets
needed to correctly train a classifier.

19
00:01:19,360 --> 00:01:23,690
We have seen data to achieve better
results and be able to generalize.

20
00:01:23,690 --> 00:01:27,870
We need to split out data set
in three independent ones.

21
00:01:27,870 --> 00:01:31,550
Training, validation, and test sets.

22
00:01:31,550 --> 00:01:36,110
Training and validation are then used
during the learning phase while the test

23
00:01:36,110 --> 00:01:42,490
set is used, is used to access the results
of a fully trained error classifier.

24
00:01:42,490 --> 00:01:47,050
And provide an, an unbiased estimate
of the generalization network.

25
00:01:47,050 --> 00:01:50,020
We have also seen

26
00:01:50,020 --> 00:01:54,940
the different cross-validation techniques
that we can use to avoid overfitting.

27
00:01:54,940 --> 00:01:57,800
And that as a quick recap
over fitting a course,

28
00:01:57,800 --> 00:02:01,190
when your model is not be
able to generalize and

29
00:02:01,190 --> 00:02:06,230
start to memorize the training data,
rather than learning the general trend.

30
00:02:06,230 --> 00:02:09,400
Like in the, in the figure on the,
on the right.

31
00:02:10,680 --> 00:02:15,060
Basically in another fitting situation
where a relatively small error on

32
00:02:15,060 --> 00:02:20,010
the training set but a much larger error
when new data is presented to the network.

33
00:02:23,080 --> 00:02:25,850
Now let's see how
the classification works.

34
00:02:25,850 --> 00:02:29,230
We can describe it as a two step process.

35
00:02:29,230 --> 00:02:32,540
First we build the model and
then we use it.

36
00:02:33,620 --> 00:02:39,180
So the first step is the learning
process in which we wish to learn

37
00:02:39,180 --> 00:02:43,670
a function that divides clearly
the determined classes.

38
00:02:43,670 --> 00:02:46,480
Typically this mapping is
sort of presented in form of

39
00:02:46,480 --> 00:02:50,070
a classification rules, decision trees,
or mathematical formulas.

40
00:02:51,910 --> 00:02:55,090
In the second step,
we want to be able to use the model to do

41
00:02:55,090 --> 00:02:57,890
the actual classification
on the real data.

42
00:02:57,890 --> 00:03:02,850
And first, we need to assess its
accuracy using an independent test set.

43
00:03:02,850 --> 00:03:05,640
And once we are satisfied
with the accuracy reached,

44
00:03:05,640 --> 00:03:13,140
we can save the classifier and deploy it
on new data whose labels are not known.

45
00:03:13,140 --> 00:03:17,050
And the outputs of a classifier
can be crispy or probabilistic.

46
00:03:18,640 --> 00:03:23,400
In a crispy classification, the output
isn't just the label so black and white.

47
00:03:23,400 --> 00:03:27,310
It's either a y or y or one y.

48
00:03:27,310 --> 00:03:29,660
IN a probabilistic classification, for

49
00:03:29,660 --> 00:03:35,410
example, the classifier, returns it's
probabilities to belong to each class.

50
00:03:36,560 --> 00:03:40,110
Probalistic classification
is very useful when, for

51
00:03:40,110 --> 00:03:44,510
example, when we need high
accuracy to make a decision.

52
00:03:44,510 --> 00:03:50,640
So we can for example take action
on the data the table probability

53
00:03:50,640 --> 00:03:56,155
to belong to a given class greater
than 90% and decided to be undecided.

54
00:03:56,155 --> 00:04:01,320
>> ...and the other cases, so,
to assign objects of two classes in

55
00:04:01,320 --> 00:04:06,960
probabilistic classification we can
apply some rules, for example the winner

56
00:04:06,960 --> 00:04:12,550
take all rule assign objects to
the class with the highest probability.

57
00:04:12,550 --> 00:04:14,860
Or we can introduce threshold and

58
00:04:14,860 --> 00:04:19,750
decide to assign objects to the class
with the highest probability but

59
00:04:19,750 --> 00:04:24,920
only if that probability is greater
than the threshold of say 40%.

60
00:04:24,920 --> 00:04:31,510
Now, let's see in more details
how to evaluate a classifier.

61
00:04:31,510 --> 00:04:37,350
As we have seen accuracy is related to
the ability of, make correct prediction.

62
00:04:38,830 --> 00:04:42,930
Speed in this case refers to
the competition of cost involved in

63
00:04:42,930 --> 00:04:45,900
building the model in the training time or

64
00:04:45,900 --> 00:04:50,640
doing the actual classification of the
data and that's the classification time.

65
00:04:52,020 --> 00:04:57,050
Robustness is the ability of the
classifier to make a correct prediction,

66
00:04:57,050 --> 00:05:00,890
in the presence of noisy or
missing values.

67
00:05:00,890 --> 00:05:05,768
Scalability, instead is the ability
to enter large amounts of data.

68
00:05:09,007 --> 00:05:10,310
Now.

69
00:05:10,310 --> 00:05:14,070
Let's briefly see some
pre-processing steps that may be

70
00:05:14,070 --> 00:05:19,100
applied to improve accuracy, efficiency,
and scalability of the classifiers.

71
00:05:19,100 --> 00:05:20,920
For example, data cleaning.

72
00:05:20,920 --> 00:05:25,660
Data cleaning refers to
the task of removing noise.

73
00:05:25,660 --> 00:05:28,770
For example,
applying the smoothing techniques.

74
00:05:28,770 --> 00:05:34,410
Or refers, or we can choose how to
treat the missing values for example

75
00:05:34,410 --> 00:05:39,890
we may decide to replace the missing
values with the most probably value.

76
00:05:41,780 --> 00:05:47,210
And actually some classification models
have already some sort of mechanism for

77
00:05:47,210 --> 00:05:49,690
handling noisy and missing data.

78
00:05:49,690 --> 00:05:53,340
But this step will still help
in achieving better results.

79
00:05:56,570 --> 00:05:59,980
Then we can do a relevance analysis,
and search, for

80
00:05:59,980 --> 00:06:05,820
example, for redundant and irrelevant
attributes that can flow down the process.

81
00:06:05,820 --> 00:06:09,850
But also lead to worse results,
because they can be misleading.

82
00:06:09,850 --> 00:06:14,500
So in this, in this type of analysis,
a correlation analysis can be

83
00:06:14,500 --> 00:06:18,810
used to then define features so
that that's statistically related.

84
00:06:18,810 --> 00:06:23,940
Basically, we want to see if two variables
change together in a consistent manner.

85
00:06:26,620 --> 00:06:30,700
For example, a strong correlation
between two attributes.

86
00:06:30,700 --> 00:06:36,610
We suggest that one of the two can be
eliminated without losing accuracy.

87
00:06:36,610 --> 00:06:39,910
Now, plots on the right
show some examples.

88
00:06:39,910 --> 00:06:43,350
Where n and e are perfect correlation.

89
00:06:43,350 --> 00:06:48,870
C and f, you can tell that
they're not correlated at all.

90
00:06:48,870 --> 00:06:54,340
D and D are still a linearly correlated
even though not perfectly and

91
00:06:54,340 --> 00:06:57,920
G is instead an example of
a non-linear correlation.

92
00:07:02,980 --> 00:07:07,970
Now let's talk more in details
about accuracy measures.

93
00:07:07,970 --> 00:07:08,940
As.

94
00:07:08,940 --> 00:07:13,500
We have already said accuracy
is better estimated on a, in a,

95
00:07:13,500 --> 00:07:15,810
on an independent test set.

96
00:07:15,810 --> 00:07:21,020
In order to avoid overoptimistic
estimates due to overfitting.

97
00:07:21,020 --> 00:07:25,870
Now, we can start defining
the classification rate

98
00:07:25,870 --> 00:07:29,638
as the overall percentage of
objects correctly classified.

99
00:07:29,638 --> 00:07:32,940
And the misclassification rate, or

100
00:07:32,940 --> 00:07:40,440
loss, is simply 1 minus
the classification rate.

101
00:07:40,440 --> 00:07:43,420
Now, the is really in visualizer.

102
00:07:43,420 --> 00:07:44,720
How good.

103
00:07:44,720 --> 00:07:48,470
Is the classification
even the class by class.

104
00:07:48,470 --> 00:07:53,360
By definition in the computation
matrix the network prediction Y

105
00:07:53,360 --> 00:07:55,830
are compared with the target T.

106
00:07:55,830 --> 00:08:00,670
Then the rows represents the true classes
and the columns the predicted classes.

107
00:08:02,820 --> 00:08:07,410
In this example we can easily
see that in the fifth.

108
00:08:07,410 --> 00:08:11,340
In the first confusion matrix
at 54 galaxies have been

109
00:08:11,340 --> 00:08:13,590
incorrectly classified as stars.

110
00:08:16,810 --> 00:08:20,320
Now let's see the confusion
matrices in more details and

111
00:08:20,320 --> 00:08:24,480
this figure shows actually
four confusion matrices so

112
00:08:24,480 --> 00:08:29,290
for training, testing,
validation, and for all the data.

113
00:08:29,290 --> 00:08:34,310
In this case we can also that
the network outputs are very accurate.

114
00:08:34,310 --> 00:08:39,300
As as you can see by the high numbers of
correct responses in the green square,

115
00:08:39,300 --> 00:08:41,130
in the green squares.

116
00:08:41,130 --> 00:08:45,280
And the low numbers of incorrect
responses in the red squares.

117
00:08:46,860 --> 00:08:50,520
Now the lower right of the squares
has showed you overall accuracy.

118
00:08:51,570 --> 00:08:56,620
And usually the classification rate
on the training is set is much

119
00:08:56,620 --> 00:09:01,040
higher than one of the test set of course
because we are using that data for the,

120
00:09:01,040 --> 00:09:02,730
in the learning process.

121
00:09:02,730 --> 00:09:07,390
But remember that it's the classification
rate on the test set that

122
00:09:07,390 --> 00:09:12,720
give you an Idea of the generalization
power of the classifier you are building.

123
00:09:14,930 --> 00:09:21,160
Now, other two important measures related
are completeness and contamination.

124
00:09:21,160 --> 00:09:24,990
Completeness can be defined as
the percentage of objects of

125
00:09:24,990 --> 00:09:29,140
a given class correctly
classified as such.

126
00:09:29,140 --> 00:09:31,650
On the other hand, contamination.

127
00:09:31,650 --> 00:09:36,870
For a given class is the percentage of
objects of other classes incorrectly

128
00:09:36,870 --> 00:09:41,980
classified as belonging to that class so
that they will contaminate that class.

129
00:09:43,040 --> 00:09:45,610
And precision can be seen as a measure,

130
00:09:45,610 --> 00:09:49,050
precision defined a one
minus the contamination.

131
00:09:49,050 --> 00:09:54,170
Can be seen as a measure of the quality of
your classification while completeness is

132
00:09:54,170 --> 00:09:56,410
more a measure of quantity.

133
00:09:56,410 --> 00:10:00,020
In simple terms high precisions means that

134
00:10:00,020 --> 00:10:04,670
the algorithm basically returns
substantially more relevant

135
00:10:04,670 --> 00:10:10,680
results that irrelevant while high
completeness means that an algorithm will.

136
00:10:10,680 --> 00:10:12,790
Retrieved most of the relevant results.

137
00:10:15,350 --> 00:10:20,300
Confusion matrix can be also
generalized for M classes.

138
00:10:20,300 --> 00:10:21,520
And in general,

139
00:10:21,520 --> 00:10:28,340
given M classes a confusion matrix is
a m by m table as shown in this example.

140
00:10:28,340 --> 00:10:31,140
And for
the classifier to have a good accuracy it

141
00:10:31,140 --> 00:10:35,870
must have the examples represented
along with the diagonal and

142
00:10:35,870 --> 00:10:40,100
the entries in the other cells
should be close to zero.

143
00:10:40,100 --> 00:10:45,650
Now in this example,
that's a supernova subtype classification.

144
00:10:45,650 --> 00:10:49,350
We can see that we are able
to classify very well

145
00:10:49,350 --> 00:10:53,140
supernova type 1a with good accuracy.

146
00:10:53,140 --> 00:11:00,160
But most of the supernova type
1B are confused with the 1A.

147
00:11:00,160 --> 00:11:03,610
So in this case the confusion
matrix can also help

148
00:11:03,610 --> 00:11:08,330
understand which classes
are overlapping in many analysis.

149
00:11:10,720 --> 00:11:14,160
Now let's consider just
a two class problem.

150
00:11:14,160 --> 00:11:17,010
For example, even though some features,

151
00:11:17,010 --> 00:11:21,640
we want to be able to tell if
a patient has or not cancer.

152
00:11:22,750 --> 00:11:26,020
Classes in this case have negative and
positive.

153
00:11:26,020 --> 00:11:30,930
Through negative and through positives
are the patients correctly classifying as

154
00:11:30,930 --> 00:11:33,120
having or not having cancer.

155
00:11:33,120 --> 00:11:37,400
False positive are the patients classified
as cancerous while they are not.

156
00:11:38,670 --> 00:11:40,760
False negative the opposite.

157
00:11:40,760 --> 00:11:45,010
Now we have defined the accuracy
rate as the percentage of

158
00:11:45,010 --> 00:11:47,710
objects correctly classified.

159
00:11:47,710 --> 00:11:54,621
Now let's ask a question in this case
an accuracy rate of 90% is satisfying.

160
00:11:58,340 --> 00:12:00,410
Actually, we cannot tell.

161
00:12:00,410 --> 00:12:06,150
No, this is a clear case where
just is not sufficient because

162
00:12:06,150 --> 00:12:12,550
suppose only 3-4% of the training
set is labelled as cancer,

163
00:12:12,550 --> 00:12:15,933
a classifier that classify
all the elements as not.

164
00:12:15,933 --> 00:12:21,460
Cancer would have a 96%
classification rate.

165
00:12:21,460 --> 00:12:26,920
And this is a common problem especially
when dealing with unbalanced data sets.

166
00:12:26,920 --> 00:12:32,350
Besides, in this case, the cost
associated with false negative will be

167
00:12:32,350 --> 00:12:35,880
far greater than the cost of
associated with false positive.

168
00:12:35,880 --> 00:12:38,670
Because incorrectly classifying.

169
00:12:38,670 --> 00:12:44,440
A cancerous patient is not cancerous will
be far more costly than the opposite,

170
00:12:44,440 --> 00:12:48,000
because the first can lead to
the patient death, the last,

171
00:12:48,000 --> 00:12:49,310
the latter just to a loss.

172
00:12:53,550 --> 00:12:55,140
But luckily.

173
00:12:55,140 --> 00:12:58,740
We can use other measures to
make better decisions and

174
00:12:58,740 --> 00:13:04,800
hopefully save the patient's life like
sensitivity, specificity, and precision.

175
00:13:04,800 --> 00:13:10,550
Sensitivity is the proportion of
positive examples correctly identified.

176
00:13:10,550 --> 00:13:14,240
Basically it's the completeness
of the positive class.

177
00:13:14,240 --> 00:13:14,960
On the other hand,

178
00:13:14,960 --> 00:13:20,550
the specificity is the proportion of
negative samples correctly identified.

179
00:13:20,550 --> 00:13:23,340
And in this case,
precision is the percentage of

180
00:13:23,340 --> 00:13:28,890
the samples labeled as cancer
that are actually cancer.

181
00:13:28,890 --> 00:13:32,252
Basically, it's the ratio
between the true positive and

182
00:13:32,252 --> 00:13:35,417
the sum of the true positive
plus the false positives.

183
00:13:39,019 --> 00:13:44,250
Now, classification also
poses many challenges.

184
00:13:44,250 --> 00:13:49,550
We will need to, probably to deal with
the multiparametric large datasets.

185
00:13:49,550 --> 00:13:52,560
Data can be sparse and heterogeneous.

186
00:13:52,560 --> 00:13:56,450
We may want to be able to perform
a reliable classification in the hash

187
00:13:56,450 --> 00:13:59,920
time with the high completeness and
the low contamination.

188
00:13:59,920 --> 00:14:04,160
And we also need to include
external knowledge in our models.

189
00:14:04,160 --> 00:14:07,210
So in the next videos we'll
see in more details some

190
00:14:07,210 --> 00:14:13,120
classification models such as neural
networks and support vector machines.

191
00:14:13,120 --> 00:14:17,060
And also how to combine different
types of classifiers in order to

192
00:14:17,060 --> 00:14:18,200
achieve better results.

