1
00:00:00,000 --> 00:00:05,209
You've seen how regularization can help
prevent over fitting. But how does it

2
00:00:05,209 --> 00:00:10,830
affect the bias and variants of a learning
algorithm? In this video, I'd like to go

3
00:00:10,830 --> 00:00:15,971
deeper into the issue of bias and
variants, and talk about how it interacts

4
00:00:15,971 --> 00:00:20,770
with, and is affected by the
regularization of your learning algorithm.

5
00:00:20,770 --> 00:00:25,842
Suppose we're fitting a high order
polynomial, like that shown here. But to

6
00:00:25,842 --> 00:00:31,258
prevent over fitting, we're going to use
regularization, by that shown here. So we

7
00:00:31,258 --> 00:00:36,149
have this regularization term to try to.
The firm. The values of the firm are

8
00:00:36,149 --> 00:00:40,925
small. And as you show the verbalization
sums from G equals one to M. Rather than

9
00:00:40,925 --> 00:00:46,303
J. Equals zero to M. Let's consider three
cases. First is the case of a very large

10
00:00:46,303 --> 00:00:51,729
value of the regularization parameter
[inaudible], such as if [inaudible] were

11
00:00:51,729 --> 00:00:57,363
equal to 10,000 it's a huge value. In this
case, all of these parameters, theta one,

12
00:00:57,363 --> 00:01:03,045
theta two, theta three, and so on will be
heavily penalized and so will end up. With

13
00:01:03,240 --> 00:01:08,838
most of these parameter values being close
to zero and the hypothesis will be roughly

14
00:01:08,838 --> 00:01:13,654
H of X just equal or approximately equal
to beta zero. So, we end up with a

15
00:01:13,654 --> 00:01:18,861
hypothesis that will more or less say
that. More or less a fat constant straight

16
00:01:18,861 --> 00:01:23,873
line. And so, this hypothesis has high
bias and the value under fits this data

17
00:01:23,873 --> 00:01:29,080
set. So, the horizontal straight line is
just not a very good model for this data

18
00:01:29,080 --> 00:01:34,222
set. At the other extreme is if we had a
very small value of launder. Such as if

19
00:01:34,222 --> 00:01:39,758
launder were equal to zero. In that case
given that were fitting a high order

20
00:01:39,758 --> 00:01:45,166
polynomial, this a usual over-fitting
setting. In that case, given that we're

21
00:01:45,166 --> 00:01:50,250
fitting a higher order [inaudible] of this
[inaudible] or with very minimal

22
00:01:50,250 --> 00:01:55,232
[inaudible], with the usual high variance
over [inaudible]. This space being

23
00:01:55,232 --> 00:02:00,227
logarithm zero, zero. We're just putting
it with our regularization. So, that it

24
00:02:00,227 --> 00:02:05,546
over fits this hypothesis. And it's only
if we have some intermediate value longer

25
00:02:05,546 --> 00:02:10,605
that is neither too large or too small.
That we end up with that parameters data

26
00:02:10,605 --> 00:02:15,844
that gives us a reasonable fit to this
data. So how can we automatically choose a

27
00:02:15,844 --> 00:02:20,613
good value for the regularization
parameter Lambda? Just to reiterate, here

28
00:02:20,613 --> 00:02:25,446
is our model, and here is our learning
algorithms objective. For the setting

29
00:02:25,446 --> 00:02:30,344
where we're using regularization, let me
define J [inaudible] of theta to be

30
00:02:30,344 --> 00:02:34,791
something different, to be the
optimization objective, but without the

31
00:02:34,791 --> 00:02:39,352
regularization term. Previously, in an
earlier video when we were not using

32
00:02:39,352 --> 00:02:44,260
regularization, I defined j train of theta
to be the same as j of theta, as a cost

33
00:02:44,260 --> 00:02:48,562
function. But when we're using
regularization when there's extra lambda

34
00:02:48,562 --> 00:02:53,289
term, we're going to define j train, my
training set error, to be just my sum of

35
00:02:53,289 --> 00:02:58,318
squared errors on the training set, or my
average squared error on the training set,

36
00:02:58,318 --> 00:03:02,759
without taking into account that
regularization term. And similarly I'm

37
00:03:02,759 --> 00:03:07,881
then also going to define the cross
validation sets error and the test set

38
00:03:07,881 --> 00:03:13,617
error as before to be the average sum of
squared errors on the cross validation and

39
00:03:13,617 --> 00:03:18,944
the test sets. So just to summarize, my
definitions of Jtrain Jcv and Jtest are

40
00:03:18,944 --> 00:03:24,271
just the average squared error or one half
of the average squared error on my

41
00:03:24,271 --> 00:03:29,393
training, validation, and test sets
without the extra regularization term. So

42
00:03:29,393 --> 00:03:34,740
this how we can automatically chose the
regularization parameter lambda. What I

43
00:03:34,740 --> 00:03:39,659
usually do is maybe have some range of
values of longer [inaudible] trials. So I

44
00:03:39,659 --> 00:03:43,716
might be considering not using
regularization. Well, here are a few

45
00:03:43,716 --> 00:03:48,942
values I might try. I might be considering
[inaudible] equals over [inaudible] and so

46
00:03:48,942 --> 00:03:53,615
on. And, you know, I usually step these up
in multiples of two. Until some, some,

47
00:03:53,615 --> 00:03:58,349
maybe larger value. If I were doing this
in multiples of two, I usually end up

48
00:03:58,349 --> 00:04:03,637
with, 10.24, it's, ten exactly. But, you
know, this is close enough. And, the, the

49
00:04:03,637 --> 00:04:08,310
third and fourth decimal places won't,
won't affect your result that much. So.

50
00:04:08,310 --> 00:04:13,342
This gives me maybe, twelve different
models that I'm trying to select amongst,

51
00:04:13,342 --> 00:04:18,503
corresponding to twelve different values
of the regularization parameters along

52
00:04:18,503 --> 00:04:23,406
there. And of course you can also go to
some values less then 0.01 or values

53
00:04:23,406 --> 00:04:28,697
larger then ten, but I've just truncated
here for convenience. During each of these

54
00:04:28,697 --> 00:04:33,858
twelve models what we can do is then the
following. We can take this first model

55
00:04:33,858 --> 00:04:39,148
with [inaudible] equals zero, and minimize
my cost [inaudible], data and this would

56
00:04:39,148 --> 00:04:44,680
give me some [inaudible] of active data.
And some. After the [inaudible] video, let

57
00:04:44,680 --> 00:04:50,754
me just denote this as theta superscript
one. Then I can take my second model with

58
00:04:50,754 --> 00:04:56,754
lambda set to 0.01 and minimize the cos
function, now using lambda equals 0.01 of

59
00:04:56,754 --> 00:05:02,976
course. If we get some different parameter
of vector theta, then we denote that theta

60
00:05:02,976 --> 00:05:08,976
two. After that I end up with theta three,
so if this is for my third model, and so

61
00:05:08,976 --> 00:05:15,198
on. And so for my final model with lambda
is set to ten [inaudible] ten or 10.24 and

62
00:05:15,198 --> 00:05:20,473
put this theta. Twelve. Next, I can take
all of these hypotheses, all of these

63
00:05:20,473 --> 00:05:26,024
parameters, and use my cross validation
set to validate them. So I can look at my,

64
00:05:26,232 --> 00:05:30,950
first model, my second model, fits with
these different values of the

65
00:05:30,950 --> 00:05:36,640
regularization parameter. And you validate
them on a cross validation set, based to

66
00:05:36,640 --> 00:05:42,261
measure the average squared error of each
of these parameter vector's data on my

67
00:05:42,261 --> 00:05:48,020
cross validation set. And, I would then
pick whichever one of these twelve models

68
00:05:48,020 --> 00:05:53,834
gives me the lowest error on the. Cross
validation set. And let's say for the sake

69
00:05:53,834 --> 00:06:00,076
of this example that I end up picking beta
five. The fifth ordered polynomial because

70
00:06:00,076 --> 00:06:05,584
that has the lowest cross validation
error. Having done that. Finally, what I

71
00:06:05,584 --> 00:06:11,679
would do if I want to report the test set
error is to take the parameter beta five.

72
00:06:11,679 --> 00:06:17,480
That, that I've selected and look at how
well it does on my test set. Once again

73
00:06:17,480 --> 00:06:23,502
here is as if it fits this parameter beta
to my cross validation set. Which is why

74
00:06:23,502 --> 00:06:28,795
I'm saving. And assign a separate test set
that I'm going to use to get a better

75
00:06:28,795 --> 00:06:34,126
estimate of how well my parameter vector
beta will generalize two previous unseen

76
00:06:34,126 --> 00:06:39,002
examples. So, that's model selection
applied to selecting the regularization

77
00:06:39,002 --> 00:06:44,397
parameter [inaudible]. The last thing I'll
like to do on this video is get a better

78
00:06:44,397 --> 00:06:49,403
understanding of how cross validation and
training error vary as we vary the

79
00:06:49,403 --> 00:06:54,539
regularization parameter [inaudible]. And
so, just a reminder, right, that was our

80
00:06:54,539 --> 00:06:59,944
original cost of JH data but for this
purpose we're going. Is to define training

81
00:06:59,944 --> 00:07:05,292
error without using regularization
parameter and cross validation error

82
00:07:05,292 --> 00:07:10,715
without using the regularization
parameter. And what I like to do is plot

83
00:07:10,715 --> 00:07:16,954
this Jtrain and plot this Jcv, meaning
just how well does my hypothesis do for on

84
00:07:16,954 --> 00:07:22,897
the training set and how well does my
hypothesis do on the cross validation set

85
00:07:22,897 --> 00:07:29,594
as I vary my regularization parameter
lambda. So as we saw earlier if lambda is

86
00:07:29,594 --> 00:07:35,848
small, then we're not using much
regularization and we run the marginal

87
00:07:35,848 --> 00:07:43,247
risk of over fitting whereas if lambda is
large that is if we were on the right part

88
00:07:43,247 --> 00:07:50,470
of this horizontal axis then with a large
value of lambda we run a higher risk of

89
00:07:50,470 --> 00:07:57,869
having a bias problem. So if you plug what
j train and jcv what you find is that for

90
00:07:57,869 --> 00:08:03,531
small values of lambda you are. You can
fit the training set relatively well

91
00:08:03,531 --> 00:08:08,806
because you're not regularizing, so for
small values of lambda the regularization

92
00:08:08,806 --> 00:08:14,147
term basically goes away and you're just
minimizing pretty much [inaudible] error.

93
00:08:14,147 --> 00:08:19,554
So when lambda is small you end up with a
small value for Jtrain, whereas if lambda

94
00:08:19,554 --> 00:08:24,960
is large then you have a high bias problem
and you might not fit your training set

95
00:08:24,960 --> 00:08:30,301
well so you end up with a value up there.
So, Jtrain of theta will tend to increase

96
00:08:30,301 --> 00:08:35,444
when lambda increases because a large
value of lambda corresponds. A high bias,

97
00:08:35,444 --> 00:08:41,357
where you might not even fit your training
set well. Whereas a small value of lambda

98
00:08:41,357 --> 00:08:46,426
corresponds to, if you take, you know,
really fit a very high degree of

99
00:08:46,426 --> 00:08:52,057
polynomial to a data, let's say. As for
the cross validation error, we end up with

100
00:08:52,057 --> 00:08:57,759
a theta like this. Where, where over here
on the right, if we have a large value of

101
00:08:57,759 --> 00:09:03,390
lambda, we may end up under fitting. And
so this is the bias regime. Whereas, and,

102
00:09:03,390 --> 00:09:09,419
and so the, cross validation error would
be high, just leave all that. [inaudible]

103
00:09:09,419 --> 00:09:14,949
JFC BF data because with high bias we
won't be fitting, we won't be doing well

104
00:09:14,949 --> 00:09:20,479
with the cross validation set. Whereas
here on the left this is the high variance

105
00:09:20,479 --> 00:09:25,531
regime. Where if we have too small of
value of launder then we may be over

106
00:09:25,531 --> 00:09:31,130
fitting the data. And so if we are over
fitting the data then the cross validation

107
00:09:31,130 --> 00:09:36,728
error will also be high. And so this is
what the, the cross validation error and

108
00:09:36,728 --> 00:09:41,575
what the training error may look like on a
training set as we vary the

109
00:09:41,575 --> 00:09:46,136
regularization. That's a launder. And, so,
once again, it will often be some

110
00:09:46,136 --> 00:09:51,037
intermediate value of launder that, you
know, [inaudible] works best in terms of

111
00:09:51,037 --> 00:09:56,001
having a small consolidation area or a
small test etc. And, where the rest of the

112
00:09:56,001 --> 00:10:00,654
crows I've drawn here are somewhat
cartoonish and somewhat idealized, so, on

113
00:10:00,654 --> 00:10:05,556
a real data set, the crows you get may
just end up looking a little bit more

114
00:10:05,556 --> 00:10:10,333
messy and just a little bit more noisy
than this. For some data sets, you will

115
00:10:10,333 --> 00:10:15,545
really see these four sort of trends and
by looking at the plot of the whole dot

116
00:10:15,545 --> 00:10:20,833
cross validation. You can either manually
or automatically try to select a point

117
00:10:20,833 --> 00:10:26,222
that minimizes the [inaudible] the cross
validation error in select the value of

118
00:10:26,222 --> 00:10:31,033
lambda corresponding to low cross
validation error. When I'm trying to pick

119
00:10:31,033 --> 00:10:35,973
the regularization parameter lambda for
learning algorithm often I find that

120
00:10:35,973 --> 00:10:41,490
plotting a figure like this one shown here
helps me understand better what's going on

121
00:10:41,490 --> 00:10:45,916
and helps me verify that I am indeed
picking a good value for the

122
00:10:45,916 --> 00:10:50,600
regularization parameter lambda. So
hopefully that gives you more insight.

123
00:10:50,600 --> 00:10:55,221
[inaudible] Who regularization. And it's
effects on the bias and gaining of the

124
00:10:55,221 --> 00:10:59,492
learning algorithm. By now, you've seen
bias and data veins from a lot of

125
00:10:59,492 --> 00:11:04,172
different perspectives. And one way to do
it in the next video is to think of all

126
00:11:04,172 --> 00:11:08,443
the insights that we've gone through and
build on them, to put together a

127
00:11:08,443 --> 00:11:13,357
diagnostic that's called learning curves,
which is a tool that I often use to try to

128
00:11:13,357 --> 00:11:17,803
diagnose with the learning algorithm.
Maybe something from a bias problem or

129
00:11:17,803 --> 00:11:20,202
variance problem or a little bit of both.
