1
00:00:00,760 --> 00:00:01,360
Hi.

2
00:00:01,360 --> 00:00:03,790
In this second module on classification,

3
00:00:03,790 --> 00:00:06,240
we'll go in more details
about the Neural Networks.

4
00:00:08,200 --> 00:00:11,890
I start giving a general
definition of Neural Network and

5
00:00:11,890 --> 00:00:14,110
focus on the Multilayer Perceptron.

6
00:00:15,240 --> 00:00:18,570
I show how to choose the different
parameters involved in

7
00:00:18,570 --> 00:00:23,329
the network building such as hidden units,
activation, and error functions.

8
00:00:25,300 --> 00:00:28,690
Artificial Neural Networks
are actually inspired,

9
00:00:28,690 --> 00:00:31,100
by the way biological
network system works.

10
00:00:32,275 --> 00:00:36,371
An Neural Networks and Artificial
Neural Network can be defined as

11
00:00:36,371 --> 00:00:41,055
a large number of highly interconnected
simple processing elements called

12
00:00:41,055 --> 00:00:44,513
neurons working together
to solve specific problems.

13
00:00:49,099 --> 00:00:55,267
Usually, a Neural Network is
structured in an input layer of neuron,

14
00:00:55,267 --> 00:00:59,456
one or more hidden layers and
one output layer.

15
00:00:59,456 --> 00:01:05,280
Neurons belonging to adjacent
layers are usually fully connected.

16
00:01:05,280 --> 00:01:09,200
And the different types of
architectures are identified both by

17
00:01:09,200 --> 00:01:13,090
the different apologies adopted for
the connections and

18
00:01:13,090 --> 00:01:17,240
by the choice of the error in the
activation functions that we'll see later.

19
00:01:18,430 --> 00:01:23,250
On this slide the left picture
represents a self organizing map,

20
00:01:23,250 --> 00:01:25,920
with all the units on one layer.

21
00:01:25,920 --> 00:01:26,770
While on the right,

22
00:01:26,770 --> 00:01:33,310
there is a multilayered perceptron with
one hidden layer with five hidden units.

23
00:01:33,310 --> 00:01:35,970
And of course, an input layer and
the output layer.

24
00:01:38,120 --> 00:01:44,000
Now, in this slide, I listed some of
the most used neural network models.

25
00:01:45,220 --> 00:01:50,150
In, in feedforwards, in feedforward
networks the connections between

26
00:01:50,150 --> 00:01:56,430
the units do not form a direct
cycle like in the first figure.

27
00:01:56,430 --> 00:02:01,780
In recurrent networks like the one
shown in the middle they do.

28
00:02:01,780 --> 00:02:06,820
Self organizing maps are typically
considered similar to feedforward but in

29
00:02:06,820 --> 00:02:12,520
reality they are quite different as with
we'll see when we talk about clustering.

30
00:02:12,520 --> 00:02:17,400
Stochastic Neural Networks are built
by introducing random aberrations into

31
00:02:17,400 --> 00:02:18,700
the network.

32
00:02:18,700 --> 00:02:22,070
And one of the most used
are the Boltzmann machine, or

33
00:02:22,070 --> 00:02:25,120
the restricted Boltzmann machine
that are shown on the right.

34
00:02:28,530 --> 00:02:33,080
Now, let's start with the basic
computational element called the note

35
00:02:33,080 --> 00:02:33,580
or unit.

36
00:02:34,610 --> 00:02:38,390
A unit receiver inputs from other units,
or

37
00:02:38,390 --> 00:02:43,030
from an external source in
the case of the input units.

38
00:02:43,030 --> 00:02:49,860
Each input has an associated weight, w,
which can be modified during the learning.

39
00:02:49,860 --> 00:02:54,560
Then, the unit computes some function
of the weighted sum of its inputs, and

40
00:02:54,560 --> 00:02:56,910
gives its, and gives it an output.

41
00:02:58,570 --> 00:03:04,000
The Multilayer Perceptron is one
of the most used supervised model.

42
00:03:04,000 --> 00:03:08,210
It consists of multiple layers
of computational units,

43
00:03:08,210 --> 00:03:11,406
usually interconnected
in a feed-forward way.

44
00:03:11,406 --> 00:03:19,260
In a Multilayer Perceptron, each neuron
in one layer has direct connections

45
00:03:19,260 --> 00:03:24,090
to all the neurons of the subsequent
layer, so it's fully interconnected.

46
00:03:27,000 --> 00:03:30,280
Now, let's see how the network learn.

47
00:03:30,280 --> 00:03:33,510
It uses a back propagation approach.

48
00:03:33,510 --> 00:03:39,640
Roughly speaking, the output of the
network is compared with the targets and

49
00:03:39,640 --> 00:03:44,720
the weights are adjusted in order
to minimize the sum of the errors.

50
00:03:47,660 --> 00:03:53,295
We have said that the network has
one input level, one output, and

51
00:03:53,295 --> 00:03:55,930
zero or more hidden layers.

52
00:03:55,930 --> 00:03:59,310
Actually, choosing that
the number of hidden layers and

53
00:03:59,310 --> 00:04:01,180
of the hidden units is crucial.

54
00:04:02,620 --> 00:04:08,450
And, and unfortunately, there is no magic
formula that we can, that we can use.

55
00:04:08,450 --> 00:04:11,950
This number depends of the,
on the number of the inputs and

56
00:04:11,950 --> 00:04:17,110
outputs, number of training cases,
the amount of noise in the targets.

57
00:04:17,110 --> 00:04:21,780
The complexity of the function to be
learned as we'll see in a moment,

58
00:04:21,780 --> 00:04:24,940
depends on the activation function and
so on.

59
00:04:24,940 --> 00:04:30,730
For example, if we have too few hidden
units they can leave the two high training

60
00:04:30,730 --> 00:04:36,630
generalization error due to
overfitting and high statistical bias.

61
00:04:36,630 --> 00:04:43,200
While if we're able too many hidden units,
we can have a low training error but

62
00:04:43,200 --> 00:04:46,280
high generalization error
because of the underfitting.

63
00:04:50,750 --> 00:04:53,990
In this slide, there are some

64
00:04:53,990 --> 00:04:59,200
possible decision boundaries which
can be generated by network.

65
00:04:59,200 --> 00:05:05,020
Adding a threshold activation function
in different numbers of hidden layers.

66
00:05:05,020 --> 00:05:10,970
For example, a the first one, is
a simple perception with just one layer.

67
00:05:10,970 --> 00:05:14,800
It can be with the linear classification.

68
00:05:14,800 --> 00:05:19,380
Two layers networks like in b,
can generate decision boundaries more

69
00:05:19,380 --> 00:05:25,420
complicated, surrounding a single
complex region of the input space.

70
00:05:25,420 --> 00:05:29,940
Networks with three or more layers
can generate arbitrary decision

71
00:05:29,940 --> 00:05:34,800
boundaries that may also be non-convex or
disjoint, like in the figure.

72
00:05:38,310 --> 00:05:42,070
Now, let's talk about activation and
derivative function.

73
00:05:42,070 --> 00:05:47,680
Changing the dysfunctions, change the
results, change not totally the results,

74
00:05:47,680 --> 00:05:51,350
but also how we can interpret the outputs.

75
00:05:51,350 --> 00:05:56,270
Activation Functions are used by most
units to transform their inputs, and

76
00:05:56,270 --> 00:06:00,490
they're also needed to introduce
non-linearity in the model.

77
00:06:00,490 --> 00:06:08,270
Some, some of the most uses are linear,
logistics, softmax and time [INAUDIBLE].

78
00:06:08,270 --> 00:06:12,950
Usually, for the output nodes, we,

79
00:06:12,950 --> 00:06:18,960
you can apply the same function as for
the heat in if the outputs are bounded.

80
00:06:18,960 --> 00:06:22,430
Otherwise, you can just supply
a linear activation function.

81
00:06:24,630 --> 00:06:29,290
In this slide, there are two
examples of activation function.

82
00:06:29,290 --> 00:06:32,230
Step function and sigmoid.

83
00:06:32,230 --> 00:06:36,170
Now, say we want to have a crispy
classification for a two class problem.

84
00:06:36,170 --> 00:06:41,490
In this case, we may use a binary
classifier with just a step activation

85
00:06:41,490 --> 00:06:47,360
function and in activate,
in, in a step function,

86
00:06:47,360 --> 00:06:52,150
the output is a certain value say,
one if the input sum is

87
00:06:52,150 --> 00:06:57,266
above a certain threshold and zero, if the
input sum is below a certain threshold.

88
00:06:57,266 --> 00:07:02,807
Sigmoid function have
the property of being similar to

89
00:07:02,807 --> 00:07:09,441
step functions but with the addition
of a region of uncertainty.

90
00:07:12,931 --> 00:07:19,093
Now, it has also been shown that the two
layers network with sigmoid hidden units

91
00:07:19,093 --> 00:07:25,813
and linear output units can approximate a
continuous function to arbitrary accuracy.

92
00:07:31,143 --> 00:07:33,873
Now, let's talk about Error Functions.

93
00:07:33,873 --> 00:07:37,864
Most methods for
supervised learning require a measure of

94
00:07:37,864 --> 00:07:43,440
the discrepancy between the network
output values and the targets.

95
00:07:43,440 --> 00:07:47,920
Now, two common uses of error functions
are the sum of the squared errors and

96
00:07:47,920 --> 00:07:48,845
the cross entropy.

97
00:07:48,845 --> 00:07:53,693
For example, using

98
00:07:53,693 --> 00:08:00,675
a Multilayer Perceptron with
a softmax activation function and

99
00:08:00,675 --> 00:08:07,140
cross-entropy error allow to interpret
the network outputs as the conditional

100
00:08:07,140 --> 00:08:12,510
probabilities p of class one given x.

101
00:08:12,510 --> 00:08:16,470
N of class two given x where
x is the input vector.

102
00:08:16,470 --> 00:08:22,190
In this way, we have an actual probability
set classification in few words

103
00:08:22,190 --> 00:08:28,150
if you want to output the probability for
each object to belong to a given class and

104
00:08:28,150 --> 00:08:35,090
you are using a Multilayer Perceptron,
you have to, use a softmax

105
00:08:35,090 --> 00:08:42,264
activation function and
a cross entropy error.

106
00:08:42,264 --> 00:08:48,636
Now, early stopping is a form of
regularization used to avoid overfitting.

107
00:08:48,636 --> 00:08:53,550
In few words, we stop the training
process when the max is,

108
00:08:53,550 --> 00:08:58,490
when the maximum number of
training cycle is reached or

109
00:08:58,490 --> 00:09:02,590
when we start to seeing an increase
in the validation error.

110
00:09:02,590 --> 00:09:05,280
That's because typically,
the validation error net

111
00:09:05,280 --> 00:09:09,210
load will initially decrease as
the model fits the data better and

112
00:09:09,210 --> 00:09:13,970
better but then start rising,
when the model begins to overfit.

113
00:09:13,970 --> 00:09:17,170
So, the idea is to stop
the learning as soon as

114
00:09:17,170 --> 00:09:19,710
the minimum in validation is achieved.

115
00:09:25,620 --> 00:09:31,130
For example, in the, in the figure,
the blue line is the training error and

116
00:09:31,130 --> 00:09:36,950
decrease always,
while the red line is valid, is

117
00:09:36,950 --> 00:09:42,950
the validation error that at some point,
reaches a minimum and starts rising.

118
00:09:42,950 --> 00:09:47,860
And with early stopping, we want to catch
that minimum and stop the training there.

119
00:09:52,020 --> 00:09:53,860
Now, in this video,

120
00:09:53,860 --> 00:09:58,620
we'll be seeing how to use a Multilayer
Perceptron for classification.

121
00:09:58,620 --> 00:10:03,530
First, of course, we need to create
a proper training validation and

122
00:10:03,530 --> 00:10:08,260
test sets and choose the cross-validation
approach we want to use.

123
00:10:08,260 --> 00:10:10,630
Then, we need to choose the topology.

124
00:10:10,630 --> 00:10:15,500
For example, the number of
hidden layers and the number of

125
00:10:15,500 --> 00:10:20,255
hidden nodes per layer and we have
said that this depend, there is not,

126
00:10:20,255 --> 00:10:27,750
that the rule of thumbs don't work in this
case and then depends on many factors.

127
00:10:27,750 --> 00:10:29,830
Then we need to choose error and

128
00:10:29,830 --> 00:10:33,990
activation functions according
to the results we want to have.

129
00:10:33,990 --> 00:10:36,702
For example, if you want to have a crispy,

130
00:10:36,702 --> 00:10:40,600
a binary crispy classifier
we can use step function and

131
00:10:40,600 --> 00:10:45,620
the sum of square errors or if you want to
have a, a probabilistic classification,

132
00:10:45,620 --> 00:10:51,754
we need to use a cross center error and
a softmax activation function.

133
00:10:51,754 --> 00:10:55,550
And then, oh, we need to choose
also the activation function for

134
00:10:55,550 --> 00:10:58,500
the output units depending on
if they are bounded or not.

135
00:10:58,500 --> 00:11:05,640
If they are bounded, we can use the same
activation function of the hidden layers.

136
00:11:05,640 --> 00:11:09,860
If we are not bounded, we can just
use a linear activation function.

137
00:11:10,950 --> 00:11:17,340
Then once we have other network,
we can run the training process.

138
00:11:17,340 --> 00:11:22,020
And assess the results using a test,
a test set, and when we

139
00:11:22,020 --> 00:11:26,990
are happy with the accuracy the precision
reached, we can freeze the network.

140
00:11:26,990 --> 00:11:31,100
And you sit on your data, so
for the actual classification.

141
00:11:32,750 --> 00:11:37,590
Now in the next,
lesson we see other models and

142
00:11:37,590 --> 00:11:41,480
those are some tools useful for
classification.

