1
00:00:01,000 --> 00:00:01,810
Hi.

2
00:00:01,810 --> 00:00:05,680
I am Ciro Donalek and in this module,
I'll be walking you through one

3
00:00:05,680 --> 00:00:09,510
of the most exciting data mining tasks,
clustering.

4
00:00:09,510 --> 00:00:12,470
We'll go over terminology, definition and

5
00:00:12,470 --> 00:00:18,050
we'll see in more details some of
the models, like and self-organizing maps.

6
00:00:18,050 --> 00:00:21,640
As in the previous ones,
this is a really broad subject.

7
00:00:21,640 --> 00:00:27,120
So the purpose of this lectures is to
give you insights on problematics of and

8
00:00:27,120 --> 00:00:29,540
also help you choosing
the right models and

9
00:00:29,540 --> 00:00:32,749
the right parameters to set and
apply in your analysis.

10
00:00:35,950 --> 00:00:39,210
Now, the difference between clustering and

11
00:00:39,210 --> 00:00:44,480
classification may not seem
great at first but let's

12
00:00:44,480 --> 00:00:49,850
remember that classification is a form of
supervised learning while clustering is.

13
00:00:49,850 --> 00:00:52,690
It's the most common unsupervised task.

14
00:00:52,690 --> 00:00:57,950
So the model is not provided with
the correct results during the training.

15
00:00:57,950 --> 00:01:00,190
According to the cluster hypothesis,

16
00:01:00,190 --> 00:01:03,570
objects in the same
cluster behave similarly.

17
00:01:03,570 --> 00:01:08,590
Meaning, that points in the same cluster
are likely to be of the same type,

18
00:01:08,590 --> 00:01:13,220
sharing for example,
some common statistical characteristics.

19
00:01:13,220 --> 00:01:19,460
With the clustering, basically we want
to find natural groupings among objects.

20
00:01:19,460 --> 00:01:25,270
So let's start with some examples,
so supposed this is our data set,

21
00:01:25,270 --> 00:01:27,450
Simpson characters.

22
00:01:27,450 --> 00:01:30,500
How would you divided them in clusters?

23
00:01:30,500 --> 00:01:33,380
And be careful because
the notion of cluster can be

24
00:01:33,380 --> 00:01:35,980
ambiguous as we see shortly.

25
00:01:38,030 --> 00:01:42,190
So, we can divide them by gender,
females and males.

26
00:01:44,350 --> 00:01:48,870
But also we could divide them
by their role in the series.

27
00:01:48,870 --> 00:01:52,659
Like Simpson family members
versus school employees.

28
00:01:55,170 --> 00:02:00,100
So these are two different but
possible clusters for this data set.

29
00:02:02,280 --> 00:02:04,450
Now, let's see another example.

30
00:02:04,450 --> 00:02:06,690
We have this set of points.

31
00:02:06,690 --> 00:02:08,120
How many clusters do you see?

32
00:02:12,586 --> 00:02:17,809
I'm pretty sure most of you would say
just two clusters because there are two

33
00:02:17,809 --> 00:02:22,150
group of points that looks,
that look really well separated.

34
00:02:24,750 --> 00:02:28,691
But some may see four clusters or

35
00:02:28,691 --> 00:02:35,970
even six little clusters,
as shown in the figure.

36
00:02:35,970 --> 00:02:42,080
From these examples, we can derive
that clustering is indeed subjective.

37
00:02:42,080 --> 00:02:46,220
But in a few moments, we will also
see some measures that we can use in

38
00:02:46,220 --> 00:02:48,330
evaluating the quality of clustering.

39
00:02:50,950 --> 00:02:55,060
Now let's first discuss
why we need clustering.

40
00:02:55,060 --> 00:03:00,950
For example, organizing data into clusters
will show the internal structure,

41
00:03:00,950 --> 00:03:03,520
and that's one of the goals
in gene clustering.

42
00:03:04,520 --> 00:03:09,960
Or we can use partition to achieve a goal
like in marketing segmentation where

43
00:03:09,960 --> 00:03:15,120
the goal is to discover distinct groups
in the customer basis and then use

44
00:03:15,120 --> 00:03:21,330
this knowledge to develop a target,
targeted advertisement, for example.

45
00:03:21,330 --> 00:03:24,030
Another task is to find anomalies or

46
00:03:24,030 --> 00:03:29,730
outliers that could be just measurement
errors or cool rare objects.

47
00:03:29,730 --> 00:03:35,680
Outliers can be defined as in the picture
where we can see two distinct clusters and

48
00:03:35,680 --> 00:03:40,620
an object that seem to not
belong to any of the clusters.

49
00:03:42,940 --> 00:03:46,980
Now in clustering of course has
many applications in almost all

50
00:03:46,980 --> 00:03:51,400
scientific fields, like astronomy,
visualization and so on.

51
00:03:54,880 --> 00:03:58,000
Now let's give a more formal definition.

52
00:03:58,000 --> 00:04:00,220
Given a set of features.

53
00:04:00,220 --> 00:04:06,140
Given a set of feature vectors D
a desired number of cluster K,

54
00:04:06,140 --> 00:04:11,890
an objective of function G, we want
to assign each feature vector to one

55
00:04:11,890 --> 00:04:17,230
cluster in order to minimize or maximize
in some cases the objective function.

56
00:04:18,880 --> 00:04:23,740
The objective function is often
defined in terms of similarity or

57
00:04:23,740 --> 00:04:28,050
distances between samples or clusters.

58
00:04:28,050 --> 00:04:31,370
Now we,
now needed to answer some basic questions.

59
00:04:31,370 --> 00:04:33,400
Like what does similar mean?

60
00:04:33,400 --> 00:04:35,640
How we can define a good partition?

61
00:04:35,640 --> 00:04:37,460
How to measure the quality and so on.

62
00:04:40,400 --> 00:04:44,260
So let's start in how
to evaluate clusters.

63
00:04:45,570 --> 00:04:49,800
We can use two criteria,
two types of criteria.

64
00:04:49,800 --> 00:04:52,720
Internal criteria and external ones.

65
00:04:52,720 --> 00:04:57,950
Internal criteria are based on
the distances between points in clusters.

66
00:04:57,950 --> 00:05:00,930
And the distances between the clusters.

67
00:05:00,930 --> 00:05:05,830
And we'll see them in a moment because
there are many ways to compute those.

68
00:05:05,830 --> 00:05:09,430
External criterion can
use direct evaluation and

69
00:05:09,430 --> 00:05:12,130
the gold standard when available.

70
00:05:12,130 --> 00:05:15,910
Direct evaluation means that experts or

71
00:05:15,910 --> 00:05:22,110
users can manually check the clusters and
apply them in the domain of interest.

72
00:05:22,110 --> 00:05:25,090
This is of course the most
direct evaluation.

73
00:05:25,090 --> 00:05:28,040
But it can be also very time consuming or

74
00:05:28,040 --> 00:05:31,400
unfeasible when we deal with
the huge amounts of data.

75
00:05:31,400 --> 00:05:32,990
So it's unfeasible to look one by one.

76
00:05:34,040 --> 00:05:38,914
Now, an alternative is to use
a labeled subset created by experts,

77
00:05:38,914 --> 00:05:44,830
where the class, where the classes are
known and compute some measures on those.

78
00:05:48,950 --> 00:05:51,610
Now let's start with
the internal measures.

79
00:05:53,120 --> 00:05:58,530
They can be divided in two big
classes: intra-cluster distances

80
00:05:58,530 --> 00:06:03,100
that are the distances between the points
belonging to the same class there.

81
00:06:03,100 --> 00:06:06,880
In the intracluster distances
that are the distances between

82
00:06:06,880 --> 00:06:10,210
clusters as shown in the figure.

83
00:06:14,150 --> 00:06:17,920
Now, typical objective
functions in clustering aim to

84
00:06:17,920 --> 00:06:21,920
obtain a low intracluster distance.

85
00:06:21,920 --> 00:06:24,530
And a high inter-cluster similarity.

86
00:06:24,530 --> 00:06:28,930
So we want all the points in
a cluster close together and

87
00:06:28,930 --> 00:06:31,800
the cluster as distant as possible.

88
00:06:33,960 --> 00:06:37,870
Now we see some of these
distances later when we talk

89
00:06:37,870 --> 00:06:43,150
about the hierarchical clustering because
they are also used to match clusters.

90
00:06:47,810 --> 00:06:52,160
Now keep in mind anyway that
good scores on internal

91
00:06:52,160 --> 00:06:58,180
criterion do not necessarily translate
in good effectiveness in an application.

92
00:06:58,180 --> 00:07:03,150
Additionally this evaluation may
be biased towards algorithms

93
00:07:03,150 --> 00:07:05,708
that use the same criteria
to build the clusters.

94
00:07:05,708 --> 00:07:10,010
For example,
[INAUDIBLE] that we see in the next video

95
00:07:10,010 --> 00:07:14,820
naturally optimize object distances,
and the distance-based internal

96
00:07:14,820 --> 00:07:20,200
criteria will likely overrate
the results of this type of clustering.

97
00:07:20,200 --> 00:07:24,980
So internal evaluations are best
suit to get some insights.

98
00:07:24,980 --> 00:07:29,280
But we can not imply that
one algorithm produce more

99
00:07:29,280 --> 00:07:33,940
valid results than another just
based on the internal criteria.

100
00:07:37,800 --> 00:07:43,210
Now, external measures instead
use a subset of labeled samples,

101
00:07:43,210 --> 00:07:45,100
when this is available.

102
00:07:45,100 --> 00:07:50,340
Keep in mind that the learning is
still unsupervised because the results

103
00:07:50,340 --> 00:07:56,580
are evaluated based on this data, that is
not used for the actual class sitting.

104
00:07:58,600 --> 00:08:05,120
And, so
this data is often created by human

105
00:08:05,120 --> 00:08:11,480
expert and
the most used measures are purity,

106
00:08:11,480 --> 00:08:17,290
normalized mutual information, rand index,
F measures, or Jaccard measures.

107
00:08:18,980 --> 00:08:20,690
Let's start with purity.

108
00:08:20,690 --> 00:08:25,120
Purity can be seen as the equivalent
of accuracy in classification.

109
00:08:25,120 --> 00:08:31,530
So we assign each cluster to
the class which is most frequent.

110
00:08:31,530 --> 00:08:34,220
And measure the purity by
counting the number of

111
00:08:34,220 --> 00:08:37,870
the correctly assigned samples per class.

112
00:08:37,870 --> 00:08:41,668
So high purity is easy to achieve
when the number of classes is large.

113
00:08:41,668 --> 00:08:47,900
In particular purity=1 if each
sample gets its own cluster and

114
00:08:47,900 --> 00:08:50,310
that not something we truly want.

115
00:08:50,310 --> 00:08:53,700
For this reason,
we cannot use purity to trade off for

116
00:08:53,700 --> 00:08:57,750
the polage of the class setting
against the number of classes.

117
00:08:57,750 --> 00:08:59,800
To do so, we need other measures.

118
00:09:02,120 --> 00:09:05,760
Now, this is an example
of bipurities computed.

119
00:09:05,760 --> 00:09:10,460
We have 3 different classes,
cross circles and diamonds.

120
00:09:10,460 --> 00:09:15,280
And let's suppose our algorithm will
find three clusters, as in the figure.

121
00:09:15,280 --> 00:09:18,880
So, the first is,
the first one is assigned to cross,

122
00:09:18,880 --> 00:09:22,020
because we find five
crosses in just one circle.

123
00:09:22,020 --> 00:09:28,720
The second is a, assigned a circle,
and the third is assigned to diamonds.

124
00:09:28,720 --> 00:09:31,870
The overall purity is 71%.

125
00:09:31,870 --> 00:09:33,620
And of course the first clu,

126
00:09:33,620 --> 00:09:37,420
the first cluster is the one
that looks more reliable.

127
00:09:37,420 --> 00:09:41,240
So we have a some sort of purity or
so cluster based.

128
00:09:45,170 --> 00:09:47,850
Now the normalized mutual information or

129
00:09:47,850 --> 00:09:53,340
NMI, measures the amount of
information by which our knowledge

130
00:09:53,340 --> 00:09:58,660
about the classes increases,
when we are told what the clusters are.

131
00:09:58,660 --> 00:10:01,990
But this measure has
the problem as purity,

132
00:10:01,990 --> 00:10:05,880
because don't penalize
large cardinalities.

133
00:10:05,880 --> 00:10:10,090
So while we usually want to
follow the rule of the data.

134
00:10:10,090 --> 00:10:13,960
Few clusters are better,
of course other things being equal.

135
00:10:19,010 --> 00:10:20,960
Now let's see other two measures.

136
00:10:20,960 --> 00:10:26,400
The Rand index measures the percentage
of decisions that are correct.

137
00:10:26,400 --> 00:10:30,640
It is defined as the,
as the sum of the true positive plus

138
00:10:30,640 --> 00:10:35,100
the true negative over
the total number of elements.

139
00:10:35,100 --> 00:10:37,780
So in this formula false positives and

140
00:10:37,780 --> 00:10:44,730
false negatives have the same weight but
as we have also seen in classification.

141
00:10:44,730 --> 00:10:49,680
We may wanted to penalize one of the two,
so we can introduce weights and

142
00:10:49,680 --> 00:10:51,020
use the F measure.

143
00:10:52,980 --> 00:10:58,970
So and if you see the formula of the F
measures there is a beta parameters and

144
00:10:58,970 --> 00:11:02,540
the if we want to penalize the false

145
00:11:02,540 --> 00:11:07,350
negative its efficient to
put beta greater than one.

146
00:11:11,210 --> 00:11:19,400
Now, lets now see how we can, how many
different types of classing there are.

147
00:11:19,400 --> 00:11:22,390
We can,
we can say that there are 4 main sets.

148
00:11:22,390 --> 00:11:25,180
And roughly speaking in
a rectangle class setting,

149
00:11:25,180 --> 00:11:30,840
we want to find a successive class
using a previously established ones.

150
00:11:30,840 --> 00:11:36,280
While in partitional, in partitional
clustering, data samples are divided in

151
00:11:36,280 --> 00:11:43,503
a non-overlapping obj, clusters.

152
00:11:43,503 --> 00:11:50,290
Then, in model based clustering, we assume
that the data are generated by a model and

153
00:11:50,290 --> 00:11:55,610
we try to recover the original
model from the data.

154
00:11:55,610 --> 00:12:02,630
The model that we recovered from
data then defines the cluster and,

155
00:12:02,630 --> 00:12:05,200
the assignment of samples to the cluster.

156
00:12:05,200 --> 00:12:07,870
So for example in the left figure,

157
00:12:07,870 --> 00:12:14,470
in the left figure we have three clusters
of points generated with the distribution.

158
00:12:14,470 --> 00:12:20,750
And we can see that the, EM algorithm
works well recovering all three of them.

159
00:12:20,750 --> 00:12:24,882
While in the density bases approach,
like the one on the right,

160
00:12:24,882 --> 00:12:32,630
classes are defined as areas of higher
density than the remainder of the dataset.

161
00:12:32,630 --> 00:12:38,560
Objects in these sparse areas are usually
considered to be noise and border points.

162
00:12:41,840 --> 00:12:45,310
Now, let's talk more about
hierarchical clustering.

163
00:12:45,310 --> 00:12:50,060
In hierarchical clustering,
we tried to find subsequent clusters using

164
00:12:50,060 --> 00:12:53,380
previously established
ones like I said before.

165
00:12:53,380 --> 00:12:56,720
Now, of course, we cannot do
an exhaustive search based on all

166
00:12:56,720 --> 00:13:01,800
possible combinations because it
would be computationally unfeasible.

167
00:13:01,800 --> 00:13:07,157
That's why we need to use some heuristics
and we can further divide the hierarchical

168
00:13:07,157 --> 00:13:12,342
cluster algorithms in agglomerative or
bottom-up and divisive or top-down.

169
00:13:12,342 --> 00:13:17,610
In a bottom-up,
approach we start with each element

170
00:13:17,610 --> 00:13:22,230
in a separate cluster and then merge
them accordingly to a given property.

171
00:13:24,010 --> 00:13:26,584
On the contrary, with a top-down approach,

172
00:13:26,584 --> 00:13:31,013
we start with all the points in just one
cluster, and then start dividing them.

173
00:13:34,540 --> 00:13:37,600
Let's see with two examples how this work.

174
00:13:37,600 --> 00:13:42,190
So, suppose we have six points
distributed as in the left figure,

175
00:13:42,190 --> 00:13:45,030
as, as in the right figure.

176
00:13:45,030 --> 00:13:48,808
So, black numbers are the points and
in red, the clusters created.

177
00:13:48,808 --> 00:13:52,210
In a bottom-up approach,
we start with four clusters.

178
00:13:53,820 --> 00:13:59,080
One in three, two in five and
four in six have their own clusters.

179
00:13:59,080 --> 00:14:01,178
And then we can start combine them.

180
00:14:06,653 --> 00:14:11,120
Now hierarchical clustering of
a course has some pro and some cons.

181
00:14:11,120 --> 00:14:15,690
In we, in hierarchical clustering,
one of the pros that this,

182
00:14:15,690 --> 00:14:19,670
that we don't need to specify,
in advance, the number of clusters.

183
00:14:19,670 --> 00:14:22,040
And we can later decide where to cut,

184
00:14:22,040 --> 00:14:24,550
in order to have the decided
number of clusters.

185
00:14:24,550 --> 00:14:30,110
So, it's also intuitive, because we are
used, we are also used to think this way.

186
00:14:30,110 --> 00:14:32,350
Unfortunately, they,
they're not scaled well.

187
00:14:33,520 --> 00:14:38,400
Like in the, like any realistic
algorithm local optima can be a problem.

188
00:14:41,280 --> 00:14:45,670
Now, so far, we have talked about
distances between clusters and

189
00:14:45,670 --> 00:14:47,460
weight to match them.

190
00:14:47,460 --> 00:14:50,800
Now we can see, detail, some alternatives.

191
00:14:50,800 --> 00:14:56,080
Single link is defined as the smallest
distance between an element in

192
00:14:56,080 --> 00:14:59,050
one cluster and an element in the other.

193
00:14:59,050 --> 00:15:04,250
This is a local measure since we pay
attention only in the area where the two

194
00:15:04,250 --> 00:15:06,630
clusters are close to each other.

195
00:15:06,630 --> 00:15:12,080
So, other more distant parts of
the clusters are not taken into account.

196
00:15:12,080 --> 00:15:15,750
So, we don't really pay attention
to the shape of the cluster when

197
00:15:15,750 --> 00:15:17,780
we use this distance.

198
00:15:18,910 --> 00:15:24,160
Now, complete link clustering takes into
account also the, takes into, into account

199
00:15:24,160 --> 00:15:29,250
the largest distance between an element in
one cluster and an element in the other.

200
00:15:30,790 --> 00:15:34,270
Basically, we measure the similarity
of two classes as the,

201
00:15:35,510 --> 00:15:38,750
is the similarity of their
most dissimilar members.

202
00:15:39,950 --> 00:15:41,940
This criterion is not local,

203
00:15:41,940 --> 00:15:47,360
because the entire structure of the
clustering can influence the decisions.

204
00:15:47,360 --> 00:15:49,510
And so this method works better for

205
00:15:49,510 --> 00:15:55,260
compact clusters with the small
diameters over long clusters.

206
00:15:55,260 --> 00:15:59,800
But its, but
it is also very sensitive to layers.

207
00:15:59,800 --> 00:16:02,140
For example, a sample, just one sample,

208
00:16:02,140 --> 00:16:09,180
far from the center can increase
the diameters of candidate merge clusters.

209
00:16:09,180 --> 00:16:11,160
And completely change
the final clustering.

210
00:16:12,410 --> 00:16:16,720
Another method that probably is the most
commonly used is the average link.

211
00:16:16,720 --> 00:16:21,680
And it's defined as the average distance
between an element in one cluster and

212
00:16:21,680 --> 00:16:23,600
an element in the other.

213
00:16:23,600 --> 00:16:29,190
Basically, the cluster quality is based
on all similarities between samples.

214
00:16:31,210 --> 00:16:34,890
And so we can avoid all the pit
force of the single link and

215
00:16:34,890 --> 00:16:36,750
complete the link criteria.

216
00:16:36,750 --> 00:16:39,610
Which equates the cluster
similarity with the similarity of

217
00:16:39,610 --> 00:16:41,720
a single pair of samples.

218
00:16:41,720 --> 00:16:46,266
Other [INAUDIBLE] involved
is the centroid medoids and

219
00:16:46,266 --> 00:16:52,690
we'll see it when we talk about
now we'll talk about similarity.

220
00:16:52,690 --> 00:16:57,730
And in generally, in general,
similarity measures are function that,

221
00:16:57,730 --> 00:17:01,970
costs are quan, are function that quantify
the similarity between the two objects.

222
00:17:01,970 --> 00:17:04,632
And they are based on distance metrics.

223
00:17:11,483 --> 00:17:15,960
Now, let's see some distance measures and
when we should use them.

224
00:17:15,960 --> 00:17:19,640
Euclidean Distances are probably
the most commonly used and

225
00:17:19,640 --> 00:17:22,280
produce a kind of
a sphere shaped clusters.

226
00:17:22,280 --> 00:17:26,840
They are often used in
optimization problems, and

227
00:17:26,840 --> 00:17:31,710
especially the square Euclidean distance,
even if it's not really a metric,

228
00:17:31,710 --> 00:17:34,730
because it doesn't satisfy
the triangle inequality.

229
00:17:35,810 --> 00:17:40,838
Now, the taxicab geometry is a form
of geometry in which the usual,

230
00:17:40,838 --> 00:17:46,170
the usual Euclidean distance ir,

231
00:17:46,170 --> 00:17:51,770
is replaced by a new metric,
in which the distance between two points.

232
00:17:51,770 --> 00:17:56,940
Is the sum of the absolute differences
of the Cartesian coordinates and

233
00:17:56,940 --> 00:18:00,360
they produce sort of
a diamond shape clusters.

234
00:18:04,430 --> 00:18:10,030
Cosine similarity is another distance
common used in information revival.

235
00:18:10,030 --> 00:18:15,020
Basically, cosine similarity should
narrate a metric that shows how

236
00:18:15,020 --> 00:18:20,590
related that to samples by looking at
the angle instead of the magnitude.

237
00:18:20,590 --> 00:18:24,660
For example, in text mining,
gives a useful

238
00:18:24,660 --> 00:18:29,420
measure of how similar two
documents are likely to be.

239
00:18:29,420 --> 00:18:35,700
In terms of their subject matter, for
example, like in the figure for similar

240
00:18:35,700 --> 00:18:41,120
scores, vectors are in the same direction
forming an angle near to zero degree.

241
00:18:42,510 --> 00:18:47,130
The mahalanobis distance is a measure
of the distance between a point and

242
00:18:47,130 --> 00:18:48,840
the distribution.

243
00:18:48,840 --> 00:18:52,430
And in the previous video, we have
also seen some other distances that

244
00:18:52,430 --> 00:18:56,110
can be used in the form of text or
binary data, like the hemming distance.

245
00:18:58,460 --> 00:19:03,870
Now, we have talked about hierarchical
clustering, let's now briefly introduce

246
00:19:03,870 --> 00:19:08,060
partitional clustering, that we'll see
more details when we speak about k.

247
00:19:09,360 --> 00:19:12,910
In partitional clustering,
we need to specify the number of

248
00:19:12,910 --> 00:19:18,350
the desired clusters k, even giving it
an input because we know what to expect or

249
00:19:18,350 --> 00:19:23,900
trying a multiple ones and
determining the best using a sum measure.

250
00:19:25,370 --> 00:19:30,050
In partitional clustering each distance
is placed the one over the clusters, and

251
00:19:30,050 --> 00:19:32,350
we can have hard and soft clustering.

252
00:19:33,530 --> 00:19:39,430
I hard clustering, each sample is
a member of one cluster exactly.

253
00:19:39,430 --> 00:19:43,500
An alternative definition about cluster,
of hard clustering is that

254
00:19:43,500 --> 00:19:49,470
a sample can be a full member of more
than one cluster, but a full member.

255
00:19:49,470 --> 00:19:54,540
So it's like a sort of crisp
classification 01, belong with a cluster,

256
00:19:54,540 --> 00:19:56,250
not belong to a cluster.

257
00:19:56,250 --> 00:20:01,480
In soft cluster each sample is
instead the degree of membership,

258
00:20:01,480 --> 00:20:04,320
of membership for each clusters.

259
00:20:04,320 --> 00:20:10,520
Some researcher, researchers also
distinguish between exhaustive clustering

260
00:20:10,520 --> 00:20:15,150
that assigns each sample to a cluster and
non-exhaustive clustering.

261
00:20:15,150 --> 00:20:21,260
When we can have some samples that
don't belong on to any of the clusters.

262
00:20:21,260 --> 00:20:25,070
And that can be useful in for
example, for outlier detection.

263
00:20:26,430 --> 00:20:31,510
Now, in summary, in this video we have
talked about the different types of

264
00:20:31,510 --> 00:20:35,892
clustering, how to choose the right model,
metrics and object,

265
00:20:35,892 --> 00:20:43,210
objective function
depending on our problem.

266
00:20:43,210 --> 00:20:46,930
Then we have introduced the hard and
soft clustering that can be

267
00:20:46,930 --> 00:20:51,930
related with the decreased probabilistic
classification for supervised algorithms.

268
00:20:51,930 --> 00:20:55,645
And we have also seen the differences
between the exhaustive and

269
00:20:55,645 --> 00:20:57,950
non-exhaustive classes.

270
00:20:57,950 --> 00:21:01,810
Now, in the next two lectures we'll
see some models in more details

