1
00:00:00,153 --> 00:00:05,917
In this video, I'd like to convey to you the
main intuitions behind how regularization works.

2
00:00:05,917 --> 00:00:10,673
We'll also write down the cost function
that we'll use when we're using regularization.

3
00:00:10,673 --> 00:00:14,271
With the hand-drawn examples that we'll have
on these slides,

4
00:00:14,271 --> 00:00:17,185
I think I'll be able to convey
part of the intuition.

5
00:00:17,185 --> 00:00:23,351
But an even better way to see for yourself how
regularization works is if you implement it

6
00:00:23,351 --> 00:00:28,333
and see it work for yourself. And if
you do the programming exercises after this,

7
00:00:28,333 --> 00:00:32,794
you get a chance to see regularization
in action for yourself.

8
00:00:32,794 --> 00:00:35,070
So, here's the intuition.

9
00:00:36,347 --> 00:00:39,989
In the previous video we saw that
if we were to fit

10
00:00:39,989 --> 00:00:44,031
a quadratic function to this data,
it gives us a pretty good fit to the data.

11
00:00:44,139 --> 00:00:49,161
Whereas if we were to fit an overly high order
degree polynomial we end up with a curve

12
00:00:49,161 --> 00:00:54,882
that may fit the training set very well,
but overfit the data poorly

13
00:00:54,882 --> 00:00:58,717
and not generalize well. Consider
the following. Suppose we were to

14
00:01:03,717 --> 00:01:05,583
penalize and make the parameters
theta 3 and theta 4 really small.

15
00:01:05,583 --> 00:01:08,135
Here's what I mean. Here's our optimization
objective, or here's our optimization

16
00:01:12,104 --> 00:01:15,035
problem, where we minimize our usual
squared error cost function.

17
00:01:17,573 --> 00:01:21,459
Let's say I take this objective, and I modify it,
and add to it, + 1000 theta 3 squared,

18
00:01:26,859 --> 00:01:29,789
+ 1000 theta 4 squared. 1000 -  I'm just
writing down is some huge number.

19
00:01:33,819 --> 00:01:37,292
Now, if we were to minimize this function,
well, the only way to make this new cost

20
00:01:37,292 --> 00:01:40,059
function small is if theta 3
and theta 4 are small.

21
00:01:44,228 --> 00:01:47,961
Because otherwise, if you have 1000 times theta 3,
this new cost function's going to be big.

22
00:01:47,961 --> 00:01:51,278
So when we minimize this new
function, we're going to end up with

23
00:01:51,278 --> 00:01:53,261
theta 3 close to zero,
and theta 4 close to zero.

24
00:01:58,953 --> 00:02:04,208
And that's as if we're getting rid of
these two terms over there. And if we do that,

25
00:02:04,208 --> 00:02:09,118
if theta 3 and theta 4 are close to zero,
then we're basically left with a quadratic

26
00:02:09,118 --> 00:02:12,267
function and so we'll end up with
a fit to the data that's a quadratic

27
00:02:12,267 --> 00:02:18,717
function plus maybe tiny contributions
from small terms, theta 3, theta 4,

28
00:02:18,717 --> 00:02:21,548
that may be very close to zero.

29
00:02:27,748 --> 00:02:31,870
So we end up with, essentially a quadratic function,
which is good, because it's a much better hypothesis.

30
00:02:35,901 --> 00:02:39,868
In this particular example, we looked at the
effect of penalizing two of the parameter

31
00:02:39,868 --> 00:02:44,344
values being large. More generally,
here's the idea behind regularization.

32
00:02:47,098 --> 00:02:51,024
The idea is that if we have small
values for the parameters,

33
00:02:55,916 --> 00:03:00,990
somehow will usually correspond
to having a simpler hypothesis.

34
00:03:00,990 --> 00:03:05,504
So for our last example, we penalize
theta 3 and theta 4, and when both

35
00:03:05,504 --> 00:03:09,455
of these were close to zero we wound up
with a much simpler hypothesis that was

36
00:03:09,455 --> 00:03:13,934
essentially a quadratic function.
But more broadly, if we penalize all

37
00:03:13,934 --> 00:03:18,471
the parameters, we can think
of that as trying to give us a simpler

38
00:03:18,471 --> 00:03:22,889
hypothesis because when these
parameters are close to zero,

39
00:03:22,889 --> 00:03:27,577
in this example that gave us a
quadratic function, but more generally,

40
00:03:28,285 --> 00:03:32,894
it's possible to show that having smaller
values of the parameters corresponds to

41
00:03:32,894 --> 00:03:38,141
usually smoother functions, thus
simpler, and which are therefore also less

42
00:03:38,141 --> 00:03:44,609
prone to overfitting. I realized that the
reasoning for why having all the parameters

43
00:03:44,609 --> 00:03:49,437
be small, why that corresponds to simpler
hypothesis, I realize that reasoning may

44
00:03:49,437 --> 00:03:53,578
not be entirely clear to you right now and
it is kind of hard to explain, unless you

45
00:03:53,578 --> 00:03:57,836
implement it yourself and see it for
yourself. But I hope that the example of

46
00:03:57,836 --> 00:04:03,150
having theta 3 and theta 4 be small,
and how that gave us a simpler hypothesis,

47
00:04:03,150 --> 00:04:07,961
I hope that helps explain why, at least
gives some intuition as to why this might

48
00:04:07,961 --> 00:04:13,952
be true. Lets look at this specific
example. For housing price prediction, we

49
00:04:13,952 --> 00:04:18,889
may have 100 features that we talked
about. Where maybe x1 is the size, x2 is

50
00:04:18,889 --> 00:04:21,744
the number of bedrooms, x3 is the
number of floors, and so on.

51
00:04:21,744 --> 00:04:27,878
And we may have 100 features. Unlike the
polynomial example, we don't know that

52
00:04:27,878 --> 00:04:32,371
theta 3, theta 4 are the high
order polynomial terms.

53
00:04:32,371 --> 00:04:38,187
So if we have just a bag, if we have just a
set of 100 features, it's hard to pick

54
00:04:38,187 --> 00:04:42,956
in advance which are the ones that are
less likely to be relevant.

55
00:04:42,956 --> 00:04:48,838
So we have 100 or 101 parameters,
and we don't know which ones to pick,

56
00:04:48,838 --> 00:04:56,011
we don't know which parameters to pick to try
to shrink. So, in regularization, what

57
00:04:56,011 --> 00:04:59,835
we're going to do is take our cost function,
here's my cost function for linear regression.

58
00:05:03,358 --> 00:05:06,299
And what I'm going to do is modify this
cost function to shrink all of my parameters.

59
00:05:06,314 --> 00:05:09,488
Because, I don't know which one or two
to try to shrink,

60
00:05:09,488 --> 00:05:13,529
so I'm going to modify my cost
function to add a term at the end.

61
00:05:18,883 --> 00:05:22,687
Like so. (And we add square brackets here as
well.) When I add an extra regularization

62
00:05:22,687 --> 00:05:28,341
term at the end to shrink every single
parameter, and so this term would tend to

63
00:05:28,341 --> 00:05:31,482
shrink all of my parameters, theta 1,
theta 2, theta 3, up to theta 100.

64
00:05:38,143 --> 00:05:41,552
By the way, by convention, the summation
here starts from one, so I'm not actually

65
00:05:41,552 --> 00:05:44,879
gonna penalize theta 0 being large, that's
a convention: that the sum is from

66
00:05:49,156 --> 00:05:53,148
I = 1 through N, rather than I = 0 through N.
But in practice it makes very little difference,

67
00:05:53,148 --> 00:05:56,774
whether you include theta 0 or not.
In practice, it will make very

68
00:05:59,667 --> 00:06:02,649
little difference to the results, but by
convention usually we regularize only

69
00:06:02,741 --> 00:06:05,165
theta 1 through theta 100.

70
00:06:07,934 --> 00:06:10,606
Writing down our regularized optimization objective,

71
00:06:10,713 --> 00:06:14,123
our regularized cost function again. Here it is.
Here's J of theta. Where this term on the right

72
00:06:14,446 --> 00:06:21,008
is a regularization term. And lamda,
here, is called the regularization parameter.

73
00:06:26,008 --> 00:06:30,143
And what lambda does, is it controls
a tradeoff between two different goals.

74
00:06:30,143 --> 00:06:34,837
The first goal captured by the first
term in the objective, is that we

75
00:06:34,837 --> 00:06:36,948
would like to fit the training data well.

76
00:06:40,687 --> 00:06:43,417
We would like to fit the training set well.
And the second goal is, we want to keep

77
00:06:43,417 --> 00:06:47,561
the parameters small, and that's captured
by the second term, by the regularization objective,

78
00:06:51,146 --> 00:06:53,366
by the regularization term.
And what lambda, the regularization

79
00:06:53,366 --> 00:06:58,765
parameter does, is it controls the trade
off between these two goals: between the

80
00:06:58,765 --> 00:07:01,980
goal of fitting the training set well, and
the goal of keeping the parameters small,

81
00:07:04,304 --> 00:07:07,426
and therefore, keeping the hypothesis
relatively simple, to avoid overfitting.

82
00:07:08,933 --> 00:07:13,872
For our housing price prediction example,
whereas previously if we had fit a very

83
00:07:14,887 --> 00:07:18,812
high order polynomial, we may have wound up
with a very wiggly or curvy function like this.

84
00:07:21,889 --> 00:07:23,402
If you still fit a high order
polynomial with all the polynomial

85
00:07:23,402 --> 00:07:27,914
features in there, but instead you
make sure to use this regularized objective.

86
00:07:31,222 --> 00:07:34,513
Then, what you can get out is,
in fact, a curve that isn't quite a

87
00:07:34,513 --> 00:07:38,304
quadratic function, but is much smoother
and much simpler. And maybe a curve like

88
00:07:40,765 --> 00:07:43,842
the magenta line that gives a much
better hypothesis for this data.

89
00:07:46,642 --> 00:07:48,722
Once again, I realize it can be
a bit difficult to see why shrinking the

90
00:07:48,722 --> 00:07:53,387
parameters can have this effect. But, if
you implement this algorithm yourself with

91
00:07:53,387 --> 00:07:56,437
regularization, you will be able
to see this effect firsthand.

92
00:08:01,714 --> 00:08:05,742
In regularized linear regression, if the
regularization parameter lambda is

93
00:08:05,742 --> 00:08:10,607
set to be very large, then what would
happen is we would end up penalizing

94
00:08:12,499 --> 00:08:16,306
the parameters theta 1, theta 2, theta 3,
theta 4 very highly.

95
00:08:19,583 --> 00:08:21,708
That is, if our hypothesis is this one down at the bottom.

96
00:08:24,370 --> 00:08:26,396
If we end up penalizing theta 1, theta 2,
theta 3, theta 4 very heavily. Then we'll end up

97
00:08:26,396 --> 00:08:30,616
with all of these parameters close to zero.
Theta 1 is close to zero. Theta 2 is close to zero.

98
00:08:34,708 --> 00:08:36,924
Theta 3 and theta 4 will end up being close to zero.

99
00:08:36,924 --> 00:08:41,223
And if we do that, it's as if we're getting rid of
these terms in our hypothesis.

100
00:08:41,223 --> 00:08:44,657
So that we're left with a hypothesis that looks like that.

101
00:08:48,457 --> 00:08:52,959
That says that housing prices are equal to theta 0,
and that is akin to fitting a flat, horizontal

102
00:08:52,959 --> 00:08:58,217
straight line to the data.
This is an example of underfitting.

103
00:09:00,186 --> 00:09:03,172
And in particular, this hypothesis, this straight line,
it fails to fit the training set well.

104
00:09:05,295 --> 00:09:07,493
It's just a fat straight line.
It doesn't go anywhere near

105
00:09:07,493 --> 00:09:10,519
most of our training examples. Another
way of saying this is that

106
00:09:13,888 --> 00:09:17,841
this hypothesis has too strong a preconception
or too high a bias that housing prices are

107
00:09:17,841 --> 00:09:21,417
equal to theta 0. Despite the
clear data to the contrary,

108
00:09:21,417 --> 00:09:25,398
it chooses to fit this flat line,
just a flat horizontal line.

109
00:09:28,383 --> 00:09:31,511
(I didn't draw that very well.)
This horizontal flat line to the data.

110
00:09:32,818 --> 00:09:39,116
So for regularization to work well, some care
should be taken to choose a good choice

111
00:09:39,116 --> 00:09:42,826
for the regularization parameter lambda.

112
00:09:46,303 --> 00:09:49,593
When we talk about multi-selection later in this
course, we'll talk about a way, a variety of ways,

113
00:09:49,623 --> 00:09:54,123
for automatically choosing the regularization
parameter lambda as well.

114
00:09:56,317 --> 00:09:59,058
So that is the idea behind regularization and
the cost function we'll use in order to use

115
00:09:59,058 --> 00:10:03,938
regularization. In the next two videos,
let's take these ideas and apply them to

116
00:10:03,938 --> 00:10:08,673
linear regression and to logistic
regression so that we can then get them to

117
00:10:08,673 --> 00:10:10,825
avoid overfitting problems.
