By now, you've seen a couple different learning algorithms: linear regression and logistic regression. They work well for many problems, but when you apply them to certain machine learning applications, they can run into a problem called overfitting, that can cause them to perform very poorly. What I'd like to do in this video is explain to you what is this overfitting problem and in the next few videos after this we'll talk about a technique called regularization, that will allow us to ameliorate or to reduce this overfitting problem and get these learning algorithms to maybe work much better. So what is overfitting? Let's keep using our running example of predicting housing prices with linear regression, where we want to predict the price as a function of the size of the house. One thing we could do is fit a linear function to this data. And if we do that, maybe we get that sort of straight line fit to the data. But this isn't a very good model. Looking at the data, it seems pretty clear that as the size of the house increases, the housing prices plateau, or flattens out as we move to the right. And so, this algorithm doesn't fit the training set very well. And we call this problem "underfitting" and another term for this is that this algorithm has "high bias". Both of these roughly mean that it's just not fitting the training data very well. The term bias is kind of historical or technical one, but the idea is that if we're fitting a straight line to the data, it's as if the algorithm has a very strong preconception or a very strong bias that housing prices are going to vary linearly with their size and despite the data to the contrary, despite the evidence to the contrary as preconceptions still are biased, still causes it to fit a straight line and just ends up being a poor fit. Now, the middle, we could fit a quadratic function to the data, and with this data set, if we fit a quadratic function maybe we get that kind of curve and that works pretty well. And at the other extreme would be if we were to fit say a fourth order polynomial to the data. Here we have five parameters theta zero through theta four. With that, we can actually build a curve that would passes through all five of our training examples. We might get a curve that looks like this. That, on the one hand, seems to do a very good job fitting the training set. And it passes through all of my data, at least. But this is a very wiggly curve. So I'm going up and down all over the place. And we don't actually think that's such a good model for predicting housing prices. This problem we call "overfitting". And another term for this is that this algorithm has high variance. The term high variance is another historical or technical one, but the intuition is that, if we're fitting such a high order polynomial, then the hypothesis can fit, it's almost as if you can fit almost any function. And the space of possible hypotheses is just too large, or it's too variable. And we don't have enough data to constrain it to give us a good hypothesis. So that's called overfitting. And in the middle, there isn't really a name. But I'm just gonna write, "just right". Where second degree polynomial, quadratic function seems to be just right for fitting this data. To recap, the problem of overfitting comes when, if we have too many features, then the learned hypothesis may fit the training set very well. So, your cost function may actually be very close to zero, maybe zero exactly. But you may then end up with a curve like this that tries too hard to fit the training set so that it fails to generalize to new examples and that fails to predict prices on new examples well. Here the term generalize refers to how well a hypothesis applies to new examples, that is to data to houses that it hasn't seen in the training set. On this slide we looked at overfitting for the case of linear regression. A similar thing can apply to logistic regression as well. Here is a logistic regression example with two features, x1 and x2. One thing we could do is fit logistic regression with just a simple hypothesis like this. Where, as usual, g is my sigmoid function. And if you do that you end up with a hypothesis trying to use maybe just a straight line to separate the positive and the negative examples. And this doesn't look like a very good fit to the hypothesis and so, once again, this is an example of underfitting or the hypothesis having high bias. In contrast, if you were to add to your features these quadratic terms then you could get a decision boundary that might look more like this. And, that's a pretty good fit to the data. Probably about as good as we could get on this training set. And finally, at the other extreme, if you were to fit a very high order polynomial, if you were to generate lots of high order polynomial terms as features, then logistic regression may contort itself, may try really hard to find a decision boundary, that fits your training data, or go to great lengths to contort itself to fit every single training example well and if the features x1 and x2 are for predicting maybe the cancer to the cancerous malignant benign beast tumors this really doesn't look like a very good hypothesis for making predictions and so once again this is an instance of overfitting and of the hypothesis having high variance and being unlikely to generalize well to new examples. Later in this course when we talk about debugging and diagnosing things that could go wrong with learning algorithms, we'll give you specific tools to recognize when overfitting, and also when underfitting may be occurring, but for now let's talk about the problem of if we think overfitting is occurring, what can we do to address it? In the previous examples we had one or two dimensional data, so we could just plot the hypothesis and see what was going on and select the appropriate degree polynomial. So earlier, for the housing prices example we can just plot the hypothesis. And see that it was fitting this very wiggly function that goes all over the place to predict housing prices. And we could then use figures like these to select an appropriate degree polynomial. So plotting the hypothesis could be one way to try to decide what degree polynomial to use, but that doesn't always work. And in fact, more often, we may have learning problems where we have a lot of features and there, it's not just a matter of selecting what the agreed polynomial. And in fact, when we have so many features, it also becomes much harder to plot the data. And it becomes much harder to visualize it, to decide what features to keep or not. So, concretely, if we're trying to predict housing prices, sometimes we can just have a lot of different features. And all of these features seem kind of useful. But, if we have a lot of features and very little training data, then overfitting can become a problem. In order to address overfitting there are two main options for things that we can do. The first option is to try to reduce the number of features. Concretely, one thing we could do is manually look through the list of features, and use that to try to decide which are the more important features. And therefore, which are the features we should keep, and which are the features we should throw out. Later in this class, we'll also talk about model selection algorithms, which are algorithms for automatically deciding which features to keep and which features to throw out. This idea of reducing the number of features can work well and can reduce overfitting and when we talk about model selection we'll go into this in much greater depth. But the disadvantage is that, by throwing away some of the features, is also throwing away some of the information you have about the problem. For example, maybe all of those features are actually useful for predicting the price of a house. So maybe we don't actually want to throw some of our information or throw some of our features away. The second option. Which we'll talk about in the next few videos is regularization. Here we're going to keep all the features, but we're going to reduce the magnitude, or the values of the parameters theta J. And this method works well, when we have a lot of features, each of which contributes a little bit to predicting the value of Y like we saw in the housing price prediction example. Or we could have a lot of features, each of which are somewhat useful, so maybe we don't want to throw them away. So this describes the idea of regularization at a very high level. And I realize that all of these details probably don't make sense to you yet. But in the next video, we'll start to formulate exactly how to apply regularization, and exactly what regularization means. And, then we'll start to figure out how to use this to make our learning algorithms work well and avoid overfitting.