Hi. My name is Ciro Donalek and this is the first of the three modules about classification. Today, I recap the basic concepts in supervised learning and then I'll show how to measure the accuracy of the classifier and how to evaluate it. In the next classes we'll see some model, some modern details. Classification is where we're [INAUDIBLE] it's one of the most popular data mining tasks. In a classification problem, we want to divide our samples in classes and assign each object to a given class. Because the class label of each training example is provided, the learning process is said to be supervised. And as we have seen in the introduction to machine learning, we talk about supervised learning when, for some examples, the correct results are known and are given in input to the model during the learning process. Ideally the classifier will be then able to generalize, that means it will be able to determine the correct results also for unseen distances. We have also talked about the data sets needed to correctly train a classifier. We have seen data to achieve better results and be able to generalize. We need to split out data set in three independent ones. Training, validation, and test sets. Training and validation are then used during the learning phase while the test set is used, is used to access the results of a fully trained error classifier. And provide an, an unbiased estimate of the generalization network. We have also seen the different cross-validation techniques that we can use to avoid overfitting. And that as a quick recap over fitting a course, when your model is not be able to generalize and start to memorize the training data, rather than learning the general trend. Like in the, in the figure on the, on the right. Basically in another fitting situation where a relatively small error on the training set but a much larger error when new data is presented to the network. Now let's see how the classification works. We can describe it as a two step process. First we build the model and then we use it. So the first step is the learning process in which we wish to learn a function that divides clearly the determined classes. Typically this mapping is sort of presented in form of a classification rules, decision trees, or mathematical formulas. In the second step, we want to be able to use the model to do the actual classification on the real data. And first, we need to assess its accuracy using an independent test set. And once we are satisfied with the accuracy reached, we can save the classifier and deploy it on new data whose labels are not known. And the outputs of a classifier can be crispy or probabilistic. In a crispy classification, the output isn't just the label so black and white. It's either a y or y or one y. IN a probabilistic classification, for example, the classifier, returns it's probabilities to belong to each class. Probalistic classification is very useful when, for example, when we need high accuracy to make a decision. So we can for example take action on the data the table probability to belong to a given class greater than 90% and decided to be undecided. >> ...and the other cases, so, to assign objects of two classes in probabilistic classification we can apply some rules, for example the winner take all rule assign objects to the class with the highest probability. Or we can introduce threshold and decide to assign objects to the class with the highest probability but only if that probability is greater than the threshold of say 40%. Now, let's see in more details how to evaluate a classifier. As we have seen accuracy is related to the ability of, make correct prediction. Speed in this case refers to the competition of cost involved in building the model in the training time or doing the actual classification of the data and that's the classification time. Robustness is the ability of the classifier to make a correct prediction, in the presence of noisy or missing values. Scalability, instead is the ability to enter large amounts of data. Now. Let's briefly see some pre-processing steps that may be applied to improve accuracy, efficiency, and scalability of the classifiers. For example, data cleaning. Data cleaning refers to the task of removing noise. For example, applying the smoothing techniques. Or refers, or we can choose how to treat the missing values for example we may decide to replace the missing values with the most probably value. And actually some classification models have already some sort of mechanism for handling noisy and missing data. But this step will still help in achieving better results. Then we can do a relevance analysis, and search, for example, for redundant and irrelevant attributes that can flow down the process. But also lead to worse results, because they can be misleading. So in this, in this type of analysis, a correlation analysis can be used to then define features so that that's statistically related. Basically, we want to see if two variables change together in a consistent manner. For example, a strong correlation between two attributes. We suggest that one of the two can be eliminated without losing accuracy. Now, plots on the right show some examples. Where n and e are perfect correlation. C and f, you can tell that they're not correlated at all. D and D are still a linearly correlated even though not perfectly and G is instead an example of a non-linear correlation. Now let's talk more in details about accuracy measures. As. We have already said accuracy is better estimated on a, in a, on an independent test set. In order to avoid overoptimistic estimates due to overfitting. Now, we can start defining the classification rate as the overall percentage of objects correctly classified. And the misclassification rate, or loss, is simply 1 minus the classification rate. Now, the is really in visualizer. How good. Is the classification even the class by class. By definition in the computation matrix the network prediction Y are compared with the target T. Then the rows represents the true classes and the columns the predicted classes. In this example we can easily see that in the fifth. In the first confusion matrix at 54 galaxies have been incorrectly classified as stars. Now let's see the confusion matrices in more details and this figure shows actually four confusion matrices so for training, testing, validation, and for all the data. In this case we can also that the network outputs are very accurate. As as you can see by the high numbers of correct responses in the green square, in the green squares. And the low numbers of incorrect responses in the red squares. Now the lower right of the squares has showed you overall accuracy. And usually the classification rate on the training is set is much higher than one of the test set of course because we are using that data for the, in the learning process. But remember that it's the classification rate on the test set that give you an Idea of the generalization power of the classifier you are building. Now, other two important measures related are completeness and contamination. Completeness can be defined as the percentage of objects of a given class correctly classified as such. On the other hand, contamination. For a given class is the percentage of objects of other classes incorrectly classified as belonging to that class so that they will contaminate that class. And precision can be seen as a measure, precision defined a one minus the contamination. Can be seen as a measure of the quality of your classification while completeness is more a measure of quantity. In simple terms high precisions means that the algorithm basically returns substantially more relevant results that irrelevant while high completeness means that an algorithm will. Retrieved most of the relevant results. Confusion matrix can be also generalized for M classes. And in general, given M classes a confusion matrix is a m by m table as shown in this example. And for the classifier to have a good accuracy it must have the examples represented along with the diagonal and the entries in the other cells should be close to zero. Now in this example, that's a supernova subtype classification. We can see that we are able to classify very well supernova type 1a with good accuracy. But most of the supernova type 1B are confused with the 1A. So in this case the confusion matrix can also help understand which classes are overlapping in many analysis. Now let's consider just a two class problem. For example, even though some features, we want to be able to tell if a patient has or not cancer. Classes in this case have negative and positive. Through negative and through positives are the patients correctly classifying as having or not having cancer. False positive are the patients classified as cancerous while they are not. False negative the opposite. Now we have defined the accuracy rate as the percentage of objects correctly classified. Now let's ask a question in this case an accuracy rate of 90% is satisfying. Actually, we cannot tell. No, this is a clear case where just is not sufficient because suppose only 3-4% of the training set is labelled as cancer, a classifier that classify all the elements as not. Cancer would have a 96% classification rate. And this is a common problem especially when dealing with unbalanced data sets. Besides, in this case, the cost associated with false negative will be far greater than the cost of associated with false positive. Because incorrectly classifying. A cancerous patient is not cancerous will be far more costly than the opposite, because the first can lead to the patient death, the last, the latter just to a loss. But luckily. We can use other measures to make better decisions and hopefully save the patient's life like sensitivity, specificity, and precision. Sensitivity is the proportion of positive examples correctly identified. Basically it's the completeness of the positive class. On the other hand, the specificity is the proportion of negative samples correctly identified. And in this case, precision is the percentage of the samples labeled as cancer that are actually cancer. Basically, it's the ratio between the true positive and the sum of the true positive plus the false positives. Now, classification also poses many challenges. We will need to, probably to deal with the multiparametric large datasets. Data can be sparse and heterogeneous. We may want to be able to perform a reliable classification in the hash time with the high completeness and the low contamination. And we also need to include external knowledge in our models. So in the next videos we'll see in more details some classification models such as neural networks and support vector machines. And also how to combine different types of classifiers in order to achieve better results.