1
00:00:00,675 --> 00:00:04,045
By now, you've seen a couple different
learning algorithms: linear regression and

2
00:00:04,045 --> 00:00:09,981
logistic regression. They work well for
many problems, but when you apply them to

3
00:00:09,981 --> 00:00:14,726
certain machine learning applications,
they can run into a problem called overfitting,

4
00:00:14,726 --> 00:00:18,848
that can cause them to perform
very poorly. What I'd like to do in this

5
00:00:18,848 --> 00:00:25,085
video is explain to you what is this overfitting
problem and in the next few videos

6
00:00:25,085 --> 00:00:29,661
after this we'll talk about a technique
called regularization, that will allow us

7
00:00:29,661 --> 00:00:34,031
to ameliorate or to reduce this overfitting
problem and get these learning

8
00:00:34,031 --> 00:00:40,772
algorithms to maybe work much better. So
what is overfitting? Let's keep using our

9
00:00:40,772 --> 00:00:46,241
running example of predicting housing
prices with linear regression, where we

10
00:00:46,241 --> 00:00:50,954
want to predict the price as a function of
the size of the house. One thing we could

11
00:00:50,954 --> 00:00:56,018
do is fit a linear function to this data.
And if we do that, maybe we get that sort

12
00:00:56,018 --> 00:01:01,111
of straight line fit to the data. But this
isn't a very good model. Looking at the

13
00:01:01,111 --> 00:01:06,011
data, it seems pretty clear that as the
size of the house increases, the housing

14
00:01:06,011 --> 00:01:11,143
prices plateau, or flattens out as
we move to the right. And so, this

15
00:01:11,143 --> 00:01:18,108
algorithm doesn't fit the training set
very well. And we call this problem "underfitting"

16
00:01:18,108 --> 00:01:24,330
and another term for this is that
this algorithm has "high bias". Both of

17
00:01:24,380 --> 00:01:30,417
these roughly mean that it's just not
fitting the training data very well. The

18
00:01:30,417 --> 00:01:35,567
term bias is kind of historical or
technical one, but the idea is that if

19
00:01:35,567 --> 00:01:39,808
we're fitting a straight line to the data,
it's as if the algorithm has a very

20
00:01:39,808 --> 00:01:46,075
strong preconception or a very strong bias
that housing prices are going to vary

21
00:01:46,075 --> 00:01:50,925
linearly with their size and despite the
data to the contrary, despite the evidence

22
00:01:50,925 --> 00:01:56,092
to the contrary as preconceptions still
are biased, still causes it to fit a

23
00:01:56,092 --> 00:02:02,653
straight line and just ends up being a
poor fit. Now, the middle, we could fit a

24
00:02:02,653 --> 00:02:07,037
quadratic function to the data, and with
this data set, if we fit a quadratic function

25
00:02:07,037 --> 00:02:11,784
maybe we get that kind of curve and that
works pretty well. And at the other

26
00:02:11,784 --> 00:02:18,094
extreme would be if we were to fit say a
fourth order polynomial to the data. Here

27
00:02:18,094 --> 00:02:22,658
we have five parameters theta zero through
theta four. With that, we can actually

28
00:02:22,658 --> 00:02:26,545
build a curve that would passes through
all five of our training examples. We

29
00:02:26,545 --> 00:02:33,695
might get a curve that looks like this.
That, on the one hand, seems to do a very

30
00:02:33,695 --> 00:02:37,450
good job fitting the training set. And
it passes through all of my data,

31
00:02:37,450 --> 00:02:40,812
at least. But this is a very
wiggly curve. So I'm going up

32
00:02:40,812 --> 00:02:44,750
and down all over the place. And we don't
actually think that's such a good model

33
00:02:44,750 --> 00:02:51,418
for predicting housing prices. This
problem we call "overfitting". And another

34
00:02:51,480 --> 00:02:59,055
term for this is that this algorithm has
high variance. The term high variance is

35
00:02:59,055 --> 00:03:04,015
another historical or technical
one, but the intuition is that, if we're

36
00:03:04,015 --> 00:03:09,414
fitting such a high order polynomial, then
the hypothesis can fit,

37
00:03:09,414 --> 00:03:12,304
it's almost as if you can fit almost any
function. And the space of possible

38
00:03:12,304 --> 00:03:17,726
hypotheses is just too large, or it's too
variable. And we don't have enough data to

39
00:03:17,726 --> 00:03:22,814
constrain it to give us a good hypothesis.
So that's called overfitting. And in the

40
00:03:22,814 --> 00:03:27,185
middle, there isn't really a name. But I'm
just gonna write, "just right".

41
00:03:27,185 --> 00:03:30,907
Where second degree polynomial, quadratic
function seems to be just right for

42
00:03:30,907 --> 00:03:35,696
fitting this data. To recap, the
problem of overfitting comes when, if we

43
00:03:40,388 --> 00:03:42,753
have too many features, then the learned
hypothesis may fit the training set very

44
00:03:42,753 --> 00:03:47,367
well. So, your cost function may actually be very
close to zero, maybe zero exactly. But you

45
00:03:52,413 --> 00:03:55,632
may then end up with a curve like this
that tries too hard to fit the

46
00:03:55,632 --> 00:03:58,958
training set so that it fails to
generalize to new examples and that fails

47
00:04:02,665 --> 00:04:06,313
to predict prices on new examples well.
Here the term generalize refers to how

48
00:04:06,313 --> 00:04:10,054
well a hypothesis applies to new
examples, that is to data to houses that

49
00:04:13,592 --> 00:04:18,665
it hasn't seen in the training set. On
this slide we looked at overfitting for

50
00:04:18,665 --> 00:04:24,035
the case of linear regression. A similar
thing can apply to logistic regression as

51
00:04:24,035 --> 00:04:26,974
well. Here is a logistic regression
example with two features, x1 and x2. One

52
00:04:31,205 --> 00:04:33,789
thing we could do is fit logistic
regression with just a simple hypothesis

53
00:04:33,789 --> 00:04:39,503
like this. Where, as usual, g is my sigmoid
function. And if you do that you end up

54
00:04:39,503 --> 00:04:44,009
with a hypothesis trying to use maybe just
a straight line to separate the positive

55
00:04:44,009 --> 00:04:47,596
and the negative examples. And this
doesn't look like a very good fit to the

56
00:04:47,596 --> 00:04:52,482
hypothesis and so, once again, this is an
example of underfitting or the hypothesis

57
00:04:52,482 --> 00:04:59,262
having high bias. In contrast, if you were
to add to your features these quadratic

58
00:04:59,262 --> 00:05:05,068
terms then you could get a decision
boundary that might look more like this.

59
00:05:05,068 --> 00:05:09,680
And, that's a pretty good fit to
the data. Probably about

60
00:05:09,680 --> 00:05:15,673
as good as we could get on this training
set. And finally, at the other extreme, if

61
00:05:15,673 --> 00:05:18,542
you were to fit a very high order
polynomial, if you were to generate lots

62
00:05:18,542 --> 00:05:23,453
of high order polynomial terms as
features, then logistic regression may

63
00:05:23,453 --> 00:05:30,171
contort itself, may try really hard to
find a decision boundary, that fits

64
00:05:30,171 --> 00:05:35,839
your training data, or go to great lengths
to contort itself to fit every single

65
00:05:35,839 --> 00:05:40,458
training example well and if the
features x1 and x2 are for predicting

66
00:05:40,458 --> 00:05:46,464
maybe the cancer to the cancerous
malignant benign beast tumors this

67
00:05:46,464 --> 00:05:51,282
really doesn't look like a
very good hypothesis for making

68
00:05:51,282 --> 00:05:55,552
predictions and so once again this is an
instance of overfitting and of the

69
00:05:55,552 --> 00:06:00,309
hypothesis having high variance and
being unlikely to generalize

70
00:06:00,309 --> 00:06:07,581
well to new examples. Later in this course
when we talk about debugging and

71
00:06:07,581 --> 00:06:11,919
diagnosing things that could go wrong with
learning algorithms, we'll give you specific

72
00:06:11,919 --> 00:06:16,381
tools to recognize when overfitting, and
also when underfitting may be occurring,

73
00:06:16,381 --> 00:06:21,801
but for now let's talk about the problem of
if we think overfitting is occurring, what

74
00:06:21,801 --> 00:06:27,615
can we do to address it? In the previous
examples we had one or two dimensional

75
00:06:27,615 --> 00:06:31,970
data, so we could just plot the
hypothesis and see what was going on and

76
00:06:31,970 --> 00:06:37,420
select the appropriate degree polynomial.
So earlier, for the housing prices example

77
00:06:37,420 --> 00:06:42,036
we can just plot the hypothesis. And see
that it was fitting this

78
00:06:42,036 --> 00:06:46,430
very wiggly function that goes all over
the place to predict housing prices. And we

79
00:06:46,430 --> 00:06:50,561
could then use figures like these to
select an appropriate degree polynomial.

80
00:06:50,561 --> 00:06:56,669
So plotting the hypothesis could be one
way to try to decide what degree

81
00:06:56,669 --> 00:07:02,099
polynomial to use, but that doesn't always
work. And in fact, more often, we may have

82
00:07:02,099 --> 00:07:07,629
learning problems where we have a lot
of features and there, it's not just a

83
00:07:07,629 --> 00:07:12,475
matter of selecting what the agreed
polynomial. And in fact, when we have

84
00:07:12,475 --> 00:07:17,678
so many features, it also becomes much
harder to plot the data. And it becomes

85
00:07:17,678 --> 00:07:21,967
much harder to visualize it, to decide
what features to keep or not.

86
00:07:21,967 --> 00:07:26,722
So, concretely, if we're trying to predict
housing prices, sometimes we can just have

87
00:07:26,722 --> 00:07:30,050
a lot of different features. And all of
these features seem

88
00:07:30,050 --> 00:07:34,939
kind of useful. But, if we have a lot
of features and very little training data,

89
00:07:34,939 --> 00:07:40,047
then overfitting can become a problem. In
order to address overfitting there are

90
00:07:40,047 --> 00:07:46,108
two main options for things that we can
do. The first option is to try to reduce

91
00:07:46,108 --> 00:07:50,991
the number of features. Concretely, one
thing we could do is manually look through

92
00:07:50,991 --> 00:07:55,358
the list of features, and use that to try
to decide which are the more important

93
00:07:55,358 --> 00:07:59,400
features. And therefore, which are the
features we should keep, and which are the

94
00:07:59,400 --> 00:08:03,692
features we should throw out. Later in
this class, we'll also talk about model

95
00:08:03,692 --> 00:08:09,146
selection algorithms, which are
algorithms for automatically deciding which

96
00:08:09,146 --> 00:08:13,883
features to keep and which features to
throw out. This idea of reducing the

97
00:08:13,883 --> 00:08:19,097
number of features can work well and can
reduce overfitting and when we talk about

98
00:08:19,097 --> 00:08:22,985
model selection we'll go into this in much
greater depth. But the disadvantage is

99
00:08:22,985 --> 00:08:27,761
that, by throwing away some of the
features, is also throwing away some of

100
00:08:27,761 --> 00:08:31,449
the information you have about the
problem. For example, maybe all of those

101
00:08:31,449 --> 00:08:35,841
features are actually useful for
predicting the price of a house. So maybe

102
00:08:35,841 --> 00:08:39,739
we don't actually want to throw some of our
information or throw some of our features

103
00:08:39,739 --> 00:08:44,287
away. The second option. Which we'll talk
about in the next few

104
00:08:46,318 --> 00:08:51,377
videos is regularization. Here we're going
to keep all the features, but we're going

105
00:08:54,177 --> 00:08:57,464
to reduce the magnitude, or the values of
the parameters theta J. And this method

106
00:09:01,633 --> 00:09:03,403
works well, when we have a lot
of features, each of which contributes

107
00:09:03,403 --> 00:09:08,928
a little bit to predicting the value of Y
like we saw in the housing

108
00:09:08,928 --> 00:09:12,710
price prediction example. Or we could have
a lot of features, each of which are

109
00:09:12,710 --> 00:09:15,706
somewhat useful, so maybe we
don't want to throw them away. So this

110
00:09:15,706 --> 00:09:24,028
describes the idea of regularization at a
very high level. And I realize that all

111
00:09:24,028 --> 00:09:28,359
of these details probably don't make sense
to you yet. But in the next video, we'll

112
00:09:28,359 --> 00:09:33,280
start to formulate exactly how to apply
regularization, and exactly what

113
00:09:33,280 --> 00:09:38,054
regularization means. And, then we'll
start to figure out how to use this to

114
00:09:38,054 --> 00:09:42,054
make our learning algorithms work well
and avoid overfitting.
