Let's conclude and kind of create an overview of what can be learned about decision trees. Decision trees are really the single most popular datamining or machine learning tool. The reason why decision trees are so, so popular is the following. First, it is very easy to understand the structure of the decision tree, right? It's even, in a sense, easy to explain to someone, why did we, why the tree made a given decision, right. We can show them a series of speeds and say, because of these conditions, this is the classification that has been made. Decision trees are also very easy to implement, both at the training stage, is not that complicated, and also once we have the tree it's very easy to use it, for classification. So, trees are also easy to use. Another thing that is important is [SOUND] that they are computationally also very cheap, in a sense that classification is easy. One caveat with, decision trees is that it's very easy to overfit. So it's very easy to create complicated trees that, that are very deep and become too intricate and don't generalize well to unseen data set, and what is also nice about decision trees is that they can do both classification, as well as regression. So, decision trees have many good, the, ex, eg, properties. Another thing that decision trees are very useful for, is what is called learning ensembles, right? Many times it turns out that it is better to build multiple decision trees, latch, let each of the individual decision trees, create its own prediction, and then, for example, use these predictions as votes, or you average them together, to make the final prediction. So, this is what it means, that you learn an ensemble of decision trees, and then you kind of take, take the individual predictions, average them up or somehow aggregate them, to make the final prediction. So one method that, that does this, is called bagging. The idea here is that we want to learn multiple trees over independent samples of the training data, and then prediction from each tree, is considered to be average predictions from all the trees that we have, and make the final prediction. It turns out that in practice this kind of bagging approaches work much better, especially if the classification or the prediction task is very hard. So, the idea is very simple, right? We will have our, initial data set, we will create, many random samples out of it. right, let's call them D star 1, D star 2, this can simply be. When we are sampling from D, we did a replacement. This will give us a new set of, training data sets, and for each of these, we'll then train a separate tree, tree 1, tree 2, tree 3, and then when a new example x comes in, we, we make each of the trees to, to give a, to give a prediction. And, and then we take the, let's say the majority vote or the average, to come out, with a final prediction. So that is basically this, the whole idea of bagging. There are a few interesting things how we could, how we could use our PLANET infrastructure that we, that we just talked about to do bagging. So let's look at this. The idea is that here, right the tree induction be, be, begin at the root, and that all, all the trees are of the bagged are pushed into MapReduce queue and then all the controller kind of has to do is to do the tree induction over, over a given, dataset of samples, right. One thing that we need to discuss when the data said this star is very huge, how do we create subsamples of it? Here basically the idea is to use, hashing. Right? So the idea is that I can take training records, and I can, use hashing, I can use multiple hash functions. And for every hash function, I, this creates me a separate, different subsets, of data, where the idea is that the records of hash into a particular range, we use them to learn the tree. right? This means that this way the sample, the same sample, is used for all nodes in a given tree. And, no, the way I define right now is that this is sampling without replacement. What we, what we want to do in bagging is to use sampling with replacement. So, the way we can do this is by using, multiple hash functions. To contin, to end our discussion of decision trees and of the machine learning topics, I want to also briefly compare the support vector machines and decision trees in which of these two methods should, methods should be used in a given case. So support vector machines are generally used for classification, meaning, usually, binary classification is the most common case. They are used for real value, valued features, right, because we think about the decision boundary in this Euclidean space. So they are not, for example, used for, for categorical features like, colors, or, weather types or something, right? SVMs are also very good when we have hundreds, or hundreds of thousands of features, right? When you have lots and lots of features. So, for example, for text, the the SVMs are really great. They also work really well when we have sparse feature sets, what this means is that, most of our feature vector or most of ours features take value 0. And, what is nice about SVMs is, that they have a very simple decision boundary. The, what we talked about is a simple linear decision boundary, which means, it's very hard to overfit. So, example applications for support vector machines are like text classifications, spam detection, exa, cases in computer vision when, where we are doing, classification and so on. On the other hand, decision trees are for much in some sense more, comp, building much more complicated decision boundaries. So, decision trees, are both used for regression and classification, having up to let's say ten different classes if we talk about classification. Decision trees can handle both real valued and categorical feature, the most serious limitation of decision trees is that they can handle, only hundreds of features, right, and only dense features, right? So, when we think about decision trees, we think that maybe we have ten, or 20, or 100 different features, and for every training example all these features are, are filled in, right? And SVM,s sorry, decision trees allow us to come up with, complicated decision boundaries, this is good. On the other hand, we have to be very careful because otherwise, we will overfit to the data, and so we need to do what is called early stopping. Are, some applications for decision trees, will for example be using will be user profile, classification, where you can have user, we describe each user with a small set of features, and now we, we would ex, want to classify these users maybe as fraudulent, or not, and this would be an example for application for decision trees. Another application would be, for example, predicting the, what is called on the web, page bounce prediction, right? Whether a person land, coming to a given page, are they just going to leave the page or are they going to make the next, click. In this case again, we could describe every person with a small feature vector, and, make the prediction.