Hi. In this second module on classification, we'll go in more details about the Neural Networks. I start giving a general definition of Neural Network and focus on the Multilayer Perceptron. I show how to choose the different parameters involved in the network building such as hidden units, activation, and error functions. Artificial Neural Networks are actually inspired, by the way biological network system works. An Neural Networks and Artificial Neural Network can be defined as a large number of highly interconnected simple processing elements called neurons working together to solve specific problems. Usually, a Neural Network is structured in an input layer of neuron, one or more hidden layers and one output layer. Neurons belonging to adjacent layers are usually fully connected. And the different types of architectures are identified both by the different apologies adopted for the connections and by the choice of the error in the activation functions that we'll see later. On this slide the left picture represents a self organizing map, with all the units on one layer. While on the right, there is a multilayered perceptron with one hidden layer with five hidden units. And of course, an input layer and the output layer. Now, in this slide, I listed some of the most used neural network models. In, in feedforwards, in feedforward networks the connections between the units do not form a direct cycle like in the first figure. In recurrent networks like the one shown in the middle they do. Self organizing maps are typically considered similar to feedforward but in reality they are quite different as with we'll see when we talk about clustering. Stochastic Neural Networks are built by introducing random aberrations into the network. And one of the most used are the Boltzmann machine, or the restricted Boltzmann machine that are shown on the right. Now, let's start with the basic computational element called the note or unit. A unit receiver inputs from other units, or from an external source in the case of the input units. Each input has an associated weight, w, which can be modified during the learning. Then, the unit computes some function of the weighted sum of its inputs, and gives its, and gives it an output. The Multilayer Perceptron is one of the most used supervised model. It consists of multiple layers of computational units, usually interconnected in a feed-forward way. In a Multilayer Perceptron, each neuron in one layer has direct connections to all the neurons of the subsequent layer, so it's fully interconnected. Now, let's see how the network learn. It uses a back propagation approach. Roughly speaking, the output of the network is compared with the targets and the weights are adjusted in order to minimize the sum of the errors. We have said that the network has one input level, one output, and zero or more hidden layers. Actually, choosing that the number of hidden layers and of the hidden units is crucial. And, and unfortunately, there is no magic formula that we can, that we can use. This number depends of the, on the number of the inputs and outputs, number of training cases, the amount of noise in the targets. The complexity of the function to be learned as we'll see in a moment, depends on the activation function and so on. For example, if we have too few hidden units they can leave the two high training generalization error due to overfitting and high statistical bias. While if we're able too many hidden units, we can have a low training error but high generalization error because of the underfitting. In this slide, there are some possible decision boundaries which can be generated by network. Adding a threshold activation function in different numbers of hidden layers. For example, a the first one, is a simple perception with just one layer. It can be with the linear classification. Two layers networks like in b, can generate decision boundaries more complicated, surrounding a single complex region of the input space. Networks with three or more layers can generate arbitrary decision boundaries that may also be non-convex or disjoint, like in the figure. Now, let's talk about activation and derivative function. Changing the dysfunctions, change the results, change not totally the results, but also how we can interpret the outputs. Activation Functions are used by most units to transform their inputs, and they're also needed to introduce non-linearity in the model. Some, some of the most uses are linear, logistics, softmax and time [INAUDIBLE]. Usually, for the output nodes, we, you can apply the same function as for the heat in if the outputs are bounded. Otherwise, you can just supply a linear activation function. In this slide, there are two examples of activation function. Step function and sigmoid. Now, say we want to have a crispy classification for a two class problem. In this case, we may use a binary classifier with just a step activation function and in activate, in, in a step function, the output is a certain value say, one if the input sum is above a certain threshold and zero, if the input sum is below a certain threshold. Sigmoid function have the property of being similar to step functions but with the addition of a region of uncertainty. Now, it has also been shown that the two layers network with sigmoid hidden units and linear output units can approximate a continuous function to arbitrary accuracy. Now, let's talk about Error Functions. Most methods for supervised learning require a measure of the discrepancy between the network output values and the targets. Now, two common uses of error functions are the sum of the squared errors and the cross entropy. For example, using a Multilayer Perceptron with a softmax activation function and cross-entropy error allow to interpret the network outputs as the conditional probabilities p of class one given x. N of class two given x where x is the input vector. In this way, we have an actual probability set classification in few words if you want to output the probability for each object to belong to a given class and you are using a Multilayer Perceptron, you have to, use a softmax activation function and a cross entropy error. Now, early stopping is a form of regularization used to avoid overfitting. In few words, we stop the training process when the max is, when the maximum number of training cycle is reached or when we start to seeing an increase in the validation error. That's because typically, the validation error net load will initially decrease as the model fits the data better and better but then start rising, when the model begins to overfit. So, the idea is to stop the learning as soon as the minimum in validation is achieved. For example, in the, in the figure, the blue line is the training error and decrease always, while the red line is valid, is the validation error that at some point, reaches a minimum and starts rising. And with early stopping, we want to catch that minimum and stop the training there. Now, in this video, we'll be seeing how to use a Multilayer Perceptron for classification. First, of course, we need to create a proper training validation and test sets and choose the cross-validation approach we want to use. Then, we need to choose the topology. For example, the number of hidden layers and the number of hidden nodes per layer and we have said that this depend, there is not, that the rule of thumbs don't work in this case and then depends on many factors. Then we need to choose error and activation functions according to the results we want to have. For example, if you want to have a crispy, a binary crispy classifier we can use step function and the sum of square errors or if you want to have a, a probabilistic classification, we need to use a cross center error and a softmax activation function. And then, oh, we need to choose also the activation function for the output units depending on if they are bounded or not. If they are bounded, we can use the same activation function of the hidden layers. If we are not bounded, we can just use a linear activation function. Then once we have other network, we can run the training process. And assess the results using a test, a test set, and when we are happy with the accuracy the precision reached, we can freeze the network. And you sit on your data, so for the actual classification. Now in the next, lesson we see other models and those are some tools useful for classification.