In this video, I'd like to convey to you the main intuitions behind how regularization works. We'll also write down the cost function that we'll use when we're using regularization. With the hand-drawn examples that we'll have on these slides, I think I'll be able to convey part of the intuition. But an even better way to see for yourself how regularization works is if you implement it and see it work for yourself. And if you do the programming exercises after this, you get a chance to see regularization in action for yourself. So, here's the intuition. In the previous video we saw that if we were to fit a quadratic function to this data, it gives us a pretty good fit to the data. Whereas if we were to fit an overly high order degree polynomial we end up with a curve that may fit the training set very well, but overfit the data poorly and not generalize well. Consider the following. Suppose we were to penalize and make the parameters theta 3 and theta 4 really small. Here's what I mean. Here's our optimization objective, or here's our optimization problem, where we minimize our usual squared error cost function. Let's say I take this objective, and I modify it, and add to it, + 1000 theta 3 squared, + 1000 theta 4 squared. 1000 - I'm just writing down is some huge number. Now, if we were to minimize this function, well, the only way to make this new cost function small is if theta 3 and theta 4 are small. Because otherwise, if you have 1000 times theta 3, this new cost function's going to be big. So when we minimize this new function, we're going to end up with theta 3 close to zero, and theta 4 close to zero. And that's as if we're getting rid of these two terms over there. And if we do that, if theta 3 and theta 4 are close to zero, then we're basically left with a quadratic function and so we'll end up with a fit to the data that's a quadratic function plus maybe tiny contributions from small terms, theta 3, theta 4, that may be very close to zero. So we end up with, essentially a quadratic function, which is good, because it's a much better hypothesis. In this particular example, we looked at the effect of penalizing two of the parameter values being large. More generally, here's the idea behind regularization. The idea is that if we have small values for the parameters, somehow will usually correspond to having a simpler hypothesis. So for our last example, we penalize theta 3 and theta 4, and when both of these were close to zero we wound up with a much simpler hypothesis that was essentially a quadratic function. But more broadly, if we penalize all the parameters, we can think of that as trying to give us a simpler hypothesis because when these parameters are close to zero, in this example that gave us a quadratic function, but more generally, it's possible to show that having smaller values of the parameters corresponds to usually smoother functions, thus simpler, and which are therefore also less prone to overfitting. I realized that the reasoning for why having all the parameters be small, why that corresponds to simpler hypothesis, I realize that reasoning may not be entirely clear to you right now and it is kind of hard to explain, unless you implement it yourself and see it for yourself. But I hope that the example of having theta 3 and theta 4 be small, and how that gave us a simpler hypothesis, I hope that helps explain why, at least gives some intuition as to why this might be true. Lets look at this specific example. For housing price prediction, we may have 100 features that we talked about. Where maybe x1 is the size, x2 is the number of bedrooms, x3 is the number of floors, and so on. And we may have 100 features. Unlike the polynomial example, we don't know that theta 3, theta 4 are the high order polynomial terms. So if we have just a bag, if we have just a set of 100 features, it's hard to pick in advance which are the ones that are less likely to be relevant. So we have 100 or 101 parameters, and we don't know which ones to pick, we don't know which parameters to pick to try to shrink. So, in regularization, what we're going to do is take our cost function, here's my cost function for linear regression. And what I'm going to do is modify this cost function to shrink all of my parameters. Because, I don't know which one or two to try to shrink, so I'm going to modify my cost function to add a term at the end. Like so. (And we add square brackets here as well.) When I add an extra regularization term at the end to shrink every single parameter, and so this term would tend to shrink all of my parameters, theta 1, theta 2, theta 3, up to theta 100. By the way, by convention, the summation here starts from one, so I'm not actually gonna penalize theta 0 being large, that's a convention: that the sum is from I = 1 through N, rather than I = 0 through N. But in practice it makes very little difference, whether you include theta 0 or not. In practice, it will make very little difference to the results, but by convention usually we regularize only theta 1 through theta 100. Writing down our regularized optimization objective, our regularized cost function again. Here it is. Here's J of theta. Where this term on the right is a regularization term. And lamda, here, is called the regularization parameter. And what lambda does, is it controls a tradeoff between two different goals. The first goal captured by the first term in the objective, is that we would like to fit the training data well. We would like to fit the training set well. And the second goal is, we want to keep the parameters small, and that's captured by the second term, by the regularization objective, by the regularization term. And what lambda, the regularization parameter does, is it controls the trade off between these two goals: between the goal of fitting the training set well, and the goal of keeping the parameters small, and therefore, keeping the hypothesis relatively simple, to avoid overfitting. For our housing price prediction example, whereas previously if we had fit a very high order polynomial, we may have wound up with a very wiggly or curvy function like this. If you still fit a high order polynomial with all the polynomial features in there, but instead you make sure to use this regularized objective. Then, what you can get out is, in fact, a curve that isn't quite a quadratic function, but is much smoother and much simpler. And maybe a curve like the magenta line that gives a much better hypothesis for this data. Once again, I realize it can be a bit difficult to see why shrinking the parameters can have this effect. But, if you implement this algorithm yourself with regularization, you will be able to see this effect firsthand. In regularized linear regression, if the regularization parameter lambda is set to be very large, then what would happen is we would end up penalizing the parameters theta 1, theta 2, theta 3, theta 4 very highly. That is, if our hypothesis is this one down at the bottom. If we end up penalizing theta 1, theta 2, theta 3, theta 4 very heavily. Then we'll end up with all of these parameters close to zero. Theta 1 is close to zero. Theta 2 is close to zero. Theta 3 and theta 4 will end up being close to zero. And if we do that, it's as if we're getting rid of these terms in our hypothesis. So that we're left with a hypothesis that looks like that. That says that housing prices are equal to theta 0, and that is akin to fitting a flat, horizontal straight line to the data. This is an example of underfitting. And in particular, this hypothesis, this straight line, it fails to fit the training set well. It's just a fat straight line. It doesn't go anywhere near most of our training examples. Another way of saying this is that this hypothesis has too strong a preconception or too high a bias that housing prices are equal to theta 0. Despite the clear data to the contrary, it chooses to fit this flat line, just a flat horizontal line. (I didn't draw that very well.) This horizontal flat line to the data. So for regularization to work well, some care should be taken to choose a good choice for the regularization parameter lambda. When we talk about multi-selection later in this course, we'll talk about a way, a variety of ways, for automatically choosing the regularization parameter lambda as well. So that is the idea behind regularization and the cost function we'll use in order to use regularization. In the next two videos, let's take these ideas and apply them to linear regression and to logistic regression so that we can then get them to avoid overfitting problems.